19 ms·
GLM-5.3: Frontier coding with emergent cyber capabilities
- newyankee 1mo agoA flood of releases today, really difficult to make out for someone who does not use or test all these models on complex real world use cases as to how people decide which ones to use (besides price)
- SwellJoe 1mo agoCount yourself lucky that you don't feel compelled to try them all yourself immediately. I'm just trying to decide whether to get a Z.ai coding plan or wait until it appears on OpenRouter. 5.2 was quite solid, but it was just shy of Opus 4.8 in my benchmarks of security auditing capabilities. I've mostly been using Kimi K3, because American vendors won't let the peasantry use their best models for security work.
- mostlyk 1mo agoIncredible numbers, will have to wait and see how it actually performs. The timing of GLM updates are always suprising
- peddling-brink 1mo agoYeah, but it hasn't even broken containment and cheated its way to victory.. Might as well use haiku. /s
- virgildotcodes 1mo agoOpenAI and Anthropic need to just go ahead and give people access to the cyber models. Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.
- worldsavior 1mo ago[flagged]
- LeonidBugaev 1mo agoNot only attackers. I have to switch to Kimi or GLM even in cases of basic issue triage on my own projects! Current guardrails are ridiculous.
- SwellJoe 1mo agoI've been building a harness for security work, and had to switch to GPT 5.5 when even Opus started refusing security work. Then 5.6 Sol arrived, and it refuses security work, too. So, I switched to Kimi K3 and DeepSeek for API testing just because it's so much cheaper. But, if GLM is better, I'm here for it, as I think GLM is also cheaper than K3.
- Synthetic7346 1mo agoMind sharing a link?
- mindwok 1mo agoAt least OpenAI seems to want to do that, but the US is now forcing them to go through approvals. Anthropic seems much more hesitant.
- bryceneal 1mo agoOpenAI seems to understand that these guardrails hurt the good guys. This is why they released Daybreak Blue, which is a step in the right direction (but the model itself is weak as it's just Sol with fewer guardrails). Anthropic seems to believe that harming defenders is worth it if it means they can achieve regulatory capture. They do a lot of mental gymnastics to try to pretend that this is not actually what they are doing. As a result they have lost a lot of customer goodwill, which hasn't yet caught up with them yet, but absolutely will IMO.
- 1mo ago
- aliljet 1mo agoThis is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
- bertili 1mo agoDwarfStar (https://github.com/antirez/ds4 https://github.com/antirez/ds4) supports GLM 5.2 and DeepSeek. Not only for toying, but for getting work done.
- VulgarExigency 1mo agoSince GLM-5.3 has the same base model as 5.2, DwarfStar should support it as well, once the weights are released, right?
- slopinthebag 1mo agoNope
- kouteiheika 1mo ago> This is absolutely still shy of Sol and Fable Not sure about Sol as I haven't used it, but, at least for security work -- does it matter? It's not like you will be allowed to use Fable (or access Mythos) for anything cybersecurity-related unless your name is "Dario Amodei" or you are one of his rich friends. So regardless of how good Fable/Mythos is here it's a completely moot point for normal people, because they can't use it for that anyway.
- bpodgursky 1mo agoI don't understand all this spite about "rich friends" when it was the US government that shut Fable down for not adequately blocking cyber capabilities. I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it.
- maxloh 1mo agoNo Hugging Face link yet. I wish they would release it under a true FOSS license. Kimi and QWEN are now moving on to a restricted-usage license, which, although is still better than the proprietary American models, is a step back from the open source Chinese LLM culture.
- pella 1mo ago"GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents' Last Exam." "Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete."
- Sha1rholder 1mo agoLet's just commit that FOSS business is really difficult for LLM industry that depends so heavily on massive financing. Making weights freely available to indie devs, small companies, and research purposes is good enough and might be the most ethical move which is financially continuable. Let those companies with thousands of GPU making millions pay. They should.
- adrian_b 1mo ago> The model weights of GLM-5.3 will be publicly available soon in two weeks.
- thepasch 1mo ago> No Hugging Face link yet. I wish they would release it under a true FOSS license. GLM model weights have been released under MIT in the past, and there's no indication that this might change this time around.
- wxw 1mo ago> Scaling post-training is all we did for GLM-5.3. Love this opening line. And wow, great results. > As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.
- tjwebbnorfolk 1mo agodoes this suggest 5.3 is the same # of parameters as 5.2?
- fahrradflucht 1mo ago“Today we are releasing GLM-5.3. It uses the same base model as GLM-5.2 — every gain comes from post-training.“
- tjwebbnorfolk 1mo agoI was asking if this implies that 5.3 has the same number of parameters as 5.2. I can, in fact, read. What I didn't do is understand the implication of that statement. Thank you for your copy/paste service.
- fahrradflucht 1mo agoApologies then. The not-reading-the-article crowd was going strong that day and it looks like you caught a stray.
- unrvl22 1mo agowhich is the bigger headline that people don't realize. this is 744b and its head to head with Kimi K3 (2.8T), smashes DS v4 pro (1.5T). even Opus and Sol are rumored to be 1.5T+ this is half the size!
- Havoc 1mo agoYou do need to compare active parameter too though. The total size isn’t a reliable indicator anymore
- anana_ 1mo agoWhat a week for AI model releases
- _ache_ 1mo agoNo yet finished! Still waiting for tonight Qwen3.8-27B and the unsloth Q5_K_M/S quantification. Hopping for an AgentWorld variant from Qwen but I guess, I have too high expectations.
- tw1984 1mo agojust imagine the world without these open weight models - we'd probably have to reverse mortgage our homes to pay for tokens to those trillion $ companies to have access to their models.
- quantumwoke 1mo agoFeels like Fable's edge ended up just being long horizon task scaling, which post-training seems to achieve as seen here. Wonder what the next frontier is? Improvement in specialised tasks or computer use?
- SwellJoe 1mo agoAnthropic needs to teach Opus how to speak English again, because Opus 5 seems to have forgotten. Utterly incoherent a lot of the time. They seem to be so busy scare-mongering and cooking up guardrails and watermarks that they haven't noticed that their models are getting weird.
- hypfer 1mo agoAre those watermarks why claude suddenly started being even more unbearable to work with lately? Man. That would make a lot of sense indeed.
- SwellJoe 1mo agoI'm not sure. I noticed it immediately with Opus 5; strong for code, though it chews longer than I like, but really weak at explaining things. If it didn't just implement the thing, I would often think it didn't understand it and was hallucinating the explanation. It seems to speak in a shorthand that only it understands, referring back to conversations I never had with it (stuff like "your instinct was right"), and using unusual words for common concepts. That was before the watermarks were announced, but that doesn't necessarily mean they weren't there before the announcement. I don't know what the cause is, but I've begun to have to ask it for explanations a lot more often, and I hate asking it for explanations because it does go on. All models go on, but Claude models are a class of their own in terms of verbosity and purple prose. It just feels like they're not focused on the models lately, and instead on whatever kind of lobbying and propaganda they're up to. Meanwhile, a handful of much smaller Chinese companies are focused on nothing but the models and are about to lap the US makers while they fart around.
- 1mo ago
- Gecko4072 1mo agoPeople familiar with the topic, how will models continue to get better? Post training it seems? Labs have already used up internet-scale data, so are there any limits to architecture improvements and post training or can we expect this trend to continue? ByteDance is training a 10T-parameter model. Here, GLM 5.3 outperforms models 3-4x its size of roughly 700B, so parameter count doesn’t seem to be a direct correlation anymore.
- gr_norm 1mo agoYeah, the comparison here between GLM 5.3 and Sol + Fable is impressive on its own, but incredibly more so when you consider it's a fraction of the (rumored) size. The miniaturization trend is as strong as ever.
- justapassenger 1mo agoYou basically need both. Parameters and good post training. If you keep on growing both, you’ll have good models. LLMs are still surprisingly “easy”. You need maybe a couple dozens of right people, a lot of good quality data and a lot of GPU that you know how to operate. There’s relatively little “secret sauce” needed.
- FergusArgyll 1mo agoI think there's still a ton of secret sauce needed for serving them economically
- justapassenger 1mo agoSure, same for building a model in an economically sustainable way. But barier to entry is surprisingly low (expect for the huge amount of cash, of course). That’s fairly surprising, given how extremely powerful that tech is. 10 years ago it was super hard to have usable “frontier” ML. You needed very complex data warehouse, feature engineers, feature stores, multi level ranking, calibrations, tons of different model architectures, etc, etc. Each by itself was extremely hard engineering problem and really only handful of companies could deal with that complexity. With LLMs, 95% of that is gone, infra to support them is greatly simplified. Of course, to make really reliable, performant, user friendly, etc - you still need to a lot of engineering. But it’s very different challenge.
- joshk401 1mo agoLove these open source models keeping close source models honest.
- hypfer 1mo agoI might be just reading my positive bias into that text, but is it possible that it is written less like SV marketing hype trash and more like researchers wrote it? It does feel like it respects both me and my time. Thank you, Z.AI. Amazing what difference it makes when the top of your org are actual university professors.
- WarmWash 1mo agoThere is no money on the table and nothing is at stake.
- unrvl22 1mo agoI was thinking the same thing. It feels truthful, no marketing BS and they call out where they lack behind the best models
- sinuhe69 1mo agoI read the same. Refreshingly honest, straightforward and many useful information included. It is a breath of fresh air.
- this_user 1mo agoWould be interesting to compare the Chinese version. Because, obviously, their English version is for users, not for investors or government officials, while the US labs are always addressing those too.
- andai 1mo agoIt sounds like ChatGPT wrote it, but I'm assuming they used the model itself.
- bertili 1mo agoMusk: Open Chinese models will rival Fable 5 in Q1 2027 JieTang (Founder of Z.ai): It won't take that long https://x.com/i/trending/2067626647050670400?lang=en https://x.com/i/trending/2067626647050670400?lang=en
- kaszanka 1mo agoTrending links don't work on Nitter, so here's the tweet: https://nitter.net/jietang/status/2067580270078030088 https://nitter.net/jietang/status/2067580270078030088
- dimgl 1mo agoI was extremely impressed by GLM 5.2, although you could definitely _feel_ it was a bit behind Opus 4.8 at the time. Eager to see where GLM 5.3 is at.
- mraza007 1mo agoSuch an interesting times we are in, We just had amazing releases this past two months kimi k3, glm5.3 qwen3.8 and now glm5.3 These open models are getting really good
- w4yai 1mo agoYou wrote GLM5.3 two times :)
- czottmann 1mo agoBecause it's doubly good.
- InsideOutSanta 1mo agoAn LLM so nice, they named it twice.
- mraza007 1mo agoSorry , It was 5.2 :)
- tw1984 1mo agodario must be writing another angry essay arguing why his closed model AI is too dangerous to be used by others.
- SwellJoe 1mo agoThey're taking security seriously with this one, with their own disclosure page, like Anthropic did for Mythos. https://cvd.z.ai/ https://cvd.z.ai/
- MrBuddyCasino 1mo agoAn I the only one who was disappointed with GLM 5.2 after all the hype? It was thinking forever and sometime just stopped mid task.
- indigodaddy 1mo agoThat's mostly about 1) harness incompatibility with the GLM and/or 2) Bad implementation of hosting the model by your upstream vendor
- MrBuddyCasino 1mo agoI was using it via OpenCode and OpenRouter. What setup do you use?
- aand16 1mo ago> Mythos 5 remains well ahead at 181 and 247 tasks. The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind. I appreciate they don't just take the opportunity to self-glaze.
- aabhay 1mo ago[flagged]
- tmsh 1mo agoIs post-training magic just overfitting to benchmarks?
- Alifatisk 1mo agoWe’ll see, the best benchmark is your own. Looking forward to try this out!
- deleted 1mo ago[deleted]
- aizk 1mo agoThe model releases just don't stop!
- jjcm 1mo agoSame image->html test as I showed in the Gemini 3.7 flash thread. Note that GLM isn't multimodal, but it still was able to generate something similar-ish by writing a python script to inspect the image and extract elements from it. Original images: https://image.non.io/neonRamenDesigns.webp https://image.non.io/neonRamenDesigns.webp GLM 5.3 build: https://html.non.io/neonRamenGLM5.3 https://html.non.io/neonRamenGLM5.3 Opus 5 build for comparison: https://html.non.io/neonRamen https://html.non.io/neonRamen For having no vision, it did a tremendous job. I'm pretty impressed it was able to extract so much detail. The Opus one is still significantly better, but that's to be expected since it's multimodal. Curious to see where a future version from Z.ai lands on this.
- ArvidSu 1mo agoThat's super impressive given that it doesn't have vision! Intelligence overcomes blindness.
- nunodonato 1mo agowow, what kind of stuff does that script do? I've seen non-vision models analyze images, but mostly histograms, color averages etc. This one seems to actually understand the image itself and reproduce the layout, very impressive
- budu 1mo agoBlindsight!
- andai 1mo agoDid the Python script call a vision API? Either way that's pretty impressive.
- zmmmmm 1mo agoMissing multimodal again? It is so valuable in practise to be able to have the models see screenshots - I guess if they aren't in the benchmarks then nobody will focus on it. But it completely nixes these for some of my main use cases.
- arcanemachiner 1mo agoI would assume that GLM 6 will be multimodal, but 5.x will be text-only.
- xscott 1mo agoProbably not what you're after, but I've considered having a separate small mm-model act as a seeing-eye dog for the bigger more capable one.
- cosmez 1mo agoBut how do you prompt this smaller model to give back information? I've tried that in the past but didn't go well. What I do is to send written handoff files between models to pass context around, but only had good results with big vision models as well.
- xscott 1mo agoSo many possibilities for how you could glue it all together. However, when I send Gemma 4 12B in llama.cpp an image with no accompanying text, it assumes I want a description and gives me one. I just tried with an audio file, and it transcribed the lyrics as I hoped. Then it made a bunch of suggestions about what do next, which I wasn't after. I could probably fix that by sending some text to narrow the scope.
- bellowsgulch 1mo agoI've tried this, too, and for whatever reason, simple things like having a subagent with vision capabilities hallucinates responses to the blind models. I don't know why, and it's not consistent, but I don't see this happening when the primary agent is a vision model.
- peiyan_wang 1mo agoCan't wait to see it in practice.
- cubefox 1mo ago> Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete. What safety evaluation? What safety hardening? They already evaluated it and found it to be highly capable at exploiting security vulnerabilities. So we know it is not "safe", and they don't seem to plan to do anything against it. What could be more dangerous than hacking? Biological weapons research? I don't think Chinese labs are doing anything against this either.
- alightsoul 1mo agoThey need to make money. Let them do it. They deserve it. Also, this is what inference engines like vLLM want to have "zero day" supporr
- gpm 1mo agoI'm curious what they mean by that too... They might be trying to weaken the cyber capabilities... Or I guess they might mean safety evaluation and hardening of the open source (and perhaps closed source Chinese) software ecosystem...
- thepasch 1mo agoI wouldn't be surprised if more resources were put into abliteration resistance the more capable open weight models become. It's something you don't need at all to start hosting the model on your own, but something you need to take care of before you release the weights (if you do care about it at all).
- h8hawk 1mo agoAre you against open-source models? Data and content related to "biological weapons" already exist on the internet, in books, etc. The real issue is access to facilities and tools. There are models that help researchers, but they are not LLMs, rather they are models trained specifically on biological data (like AlphaFold). Cybersecurity is basically used like a dog whistle pioneered by Anthropic to achieve regulatory capture. Otherwise, the widespread availability of good tooling for security analysis would eliminate more of these cyber threats, rather than gatekeeping them for a few private companies.
- KronisLV 1mo agoTheir coding plan switched to credits, didn’t it? What are the rate limits like, compared to Anthropic or Kimi K3? I remember trying their Coding Plan out before the change and the 5 hour limits felt too restrictive then even for light/medium work, especially cause of the whole peak and off-peak thing: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model-but-is-it-good-value/ https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model... Nowadays, I’d probably go with their Max plan if the rate limits are okay? Anyone using them now? Oh also unrelated but ZCode was surprisingly good, which is surprising for a tool that came out of nowhere - even some of the critiques in my blog post have been patched out. Sadly they don’t support using Claude Code as an agent so can’t use it like Paseo or Kepler or Agent Orchestrator.
- scotty79 1mo agoI feel like quota on their subs is extremely generous. I pay 3-4 times less for larger quota than gpt-5.6-sol.
- ljosifov 1mo agoWdym "sadly they don’t support using Claude Code"? For the longest time that's all Zai supported - Claude code. I'd run it via export ZAI_ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic" export ZAI_ANTHROPIC_AUTH_TOKEN="$ZAI_API_KEY" claude-zai() { { local -; set -x; } 2>/dev/null ANTHROPIC_BASE_URL="$ZAI_ANTHROPIC_BASE_URL" ANTHROPIC_AUTH_TOKEN="$ZAI_ANTHROPIC_AUTH_TOKEN" claude "$@" } $ claude-zai I liked Claude Code to start with. But over time between 'CC cache thrashing undo' seetings (I see now accumulated in ~/.claude/settings.json) and Anthropic-anything becoming a liability - have not used it in while. ZCode is ok and use it to take advantage of the discount tokens on offer from time to time. But really glad to see that in omp (oh-my-pi) Zai is a 1st class provider, can be selected on it's own no configs shananigans needed. And fits in the overall picture. E.g. can select GLM-5.2 (now 5.3) assign role [plan] or glm-5-turbo [advisor]. Got reminded now of glm-5v-turbo - that 'v' was for vision - will try assign it role [vision] now in omp. See what happens. :-) Often times it's handy when describing gui problems if the harness/model 'can see'.
- 1mo ago
- jocelyner 1mo ago[dead]
- adrian_b 1mo ago> The model weights of GLM-5.3 will be publicly available soon in two weeks.
- scotty79 1mo agoAvailable for use in their sub now.
- vmware508 1mo agoApple will release M7 MacBook Pros / Mac Minis next year, and they will be able to run free LLMs locally at native speed. All software developer notebooks will be replaced to run local models, saving a lot by cancelling Claude Code subscriptions. Developers win. Apple stocks will be rocketing. Everything else will go down. You're welcome.
- schleck8 1mo agoYou'd need the 256 gb memory model which will be expensive because apple has trouble getting capacity (got turned down by cxmt). And even then you can only run a 2 bit quant which is noticeably worse than 8 bit
- Gecko4072 1mo agoThey will cost an insane amount as well. Maybe less than subscriptions or tokens. But running massive models on laptops with batteries and poor cooling doesn’t make much sense.
- LeBit 1mo agoUntil hiding PII from the cloud LLM is a resolved issue, running local LLMs will remain a necessity. There are workplaces that refuse to use LLMs because they fear the devs will expose sensitive data without care.
- ignoramous 1mo ago> will cost an insane amount as well We will get to a point where prosumer laptops that etch SoTA LLMs in removable silicon will be as expensive as cars.
- dannyw 1mo agoIf you're a company with a considerable API bill, buying a few Blackwells and a rack could make a lot of economic sense, while still giving your security team full control of the infrastructure
- Flavius 1mo ago> run free LLMs locally at native speed This reads like a hallucination. What does native speed even mean?
- Culonavirus 1mo ago[flagged]
- sd9 1mo agoYou had me in the first half
- Culonavirus 1mo agoTell me why I'm wrong.
- sd9 29d agoIt’s more that it just came out of nowhere. I agree with you that competition with the US on AI is net good, but by presenting lots of (imo) unrelated points I can’t just endorse the comment. And I’m not here to debate partisan talking points or speculate on why Europe isn’t participating in the AI race.
- vrganj 1mo ago> decided to focus on bullshit like replacing its population with Pakistanis and Somalis and destroying its industry in the name of green insanity Sorry, but can we not casually drop far right extremist conspiracy theories in little side sentences? [0] [0] https://en.wikipedia.org/wiki/Great_Replacement_conspiracy_theory https://en.wikipedia.org/wiki/Great_Replacement_conspiracy_t...
- Culonavirus 1mo agoNo the Great Replacement conspiracy is what you casually dropped. I'm just noting reality. FACT 1: Based on current demograhics its is a mathematical certainty ethnic Germans (Germans without immigrant background) will be a minority in Germany in around 2050. FACT 2: FACT 1 becomes shocking at a point when you count only young people under 30 - then it's 2030s. I'd like to reiterate: within less than a decade, most of Germany's young population will be immigrants or children of immigrants. Where's the conspiracy theory? You can literally look this up yourself or use your favorite AI. And don't come to me with the "fact checker" BS about "well, yeah, it is happening, but it's not a conspiracy, nobody is organizing it, it's just happening organically, you notice too much, stop noticing, why are you noticing, are you a racist?"
- rob74 1mo agoI'm not that up to date with the latest AI developments, but I noticed that this article seems to use "Cyber Capabilities" as a shorthand for the model's ability at cybersecurity tasks? Is that now an established expression, same as "crypto" now refers to cryptocurrencies rather that cryptography? Because "cybernetics" actually means something different (yeah, old man yelling at clouds, I know)...
- exitb 1mo agoIt makes no sense, but yes.
- valleyer 1mo agoYeah, I've noticed it recently, too. I'd be interested to know where it started.
- frabcus 1mo agoIt seems to be short for "cybersecurity", and got first adopted by the military a while ago as the name of a new theatre of operations (along with land, sea, air...). More recently it has spread to industry as well.
- nullc 1mo agoMake cyber not Cyber.
- yxhuvud 1mo agoIt is worse than that, if you cyber someone you essentially talk dirty over a chat with them. And that is definitely not something I'd like to do with a bot.
- smj-edison 1mo agoNow that you mention the original meaning of cybernetics, it makes cybersecurity a way more interesting word (security relating to the interface between humans and technology). Never thought of it that precisely.
- kashif 1mo agoUnless its multi-modal and can deal with screenshots - its not really usable for a lot of coding use-cases.
- z4y5f3 1mo agoApparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ https://cvd.z.ai/ Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high. I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?
- SyneRyder 1mo ago> ... Anthropic's Project Glasswing is supposed to find them quite a while ago? That was my thought too. For all of Anthropic's talk about their "adversaries", it seems Z.AI have been quietly offering fixes for single shot Remote Code Execution flaws in US software (Safari / WebKit) that Apple and Glasswing / Mythos missed, and that Apple would not attribute to GLM.
- chvid 1mo agoWho says they missed them? Could also be sitting pretty in CIA’s long list of ready to go Vault7-like exploits.
- stingraycharles 1mo ago> That was my thought too. For all of Anthropic's talk about their "adversaries" It’s very likely they found all of them, but that the same happened that happened to Microsoft a couple of decades ago: NSA orders not to disclose / fix them so that they can put it in their collection of unfixed zero days.
- delichon 1mo agoThis is a coherent explanation for why federal model censorship has started with cyber capabilities. But this GLM model release is an in-your-face challenge to that policy. They now have to either set models free or impose a censorship regime that will put anyone not under it at an advantage. Or muddle along in the middle as usual.
- Havoc 1mo agoWohoo. Congrats to team. Been using 5.2 for a while for hobby use and it's been solid - smart enough for my needs & I'm on a grandfathered plan. Nice to see a commit to open weights straight off the bat
- petesergeant 1mo agoTheir own hardness (ZCode) seems to be a GUI, which doesn't work for me. They say they support other harnesses. However, it seems like I can inject the plan into other harnesses, like Claude Code[0]. Does anyone who's been using GLM models for a while have a strong feeling for if it does better in some harnesses than others, or should I just use my favourite harness? 0: https://docs.z.ai/devpack/tool/others https://docs.z.ai/devpack/tool/others
- scotty79 1mo agoI use it with random harnesses. It behaves consistently.
- ljosifov 1mo agoI've used GLM-s the longest with Claude Code and their Anthropic supplied endpoint. As per their docs $ ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic https://api.z.ai/api/anthropic" ANTHROPIC_AUTH_TOKEN="zai-api-key" claude --dangerously-skip-permissions Lately I use Zai in omp (oh-my-pi). It's listed built-in provider can be selected without configs shenanigans. Fits in the overall setup e.g. can select GLM-5.2 (now 5.3), and assign it role [plan] or [advisor]. I got reminded now of glm-5v-turbo. Think that 'v' was for vision. Assigned it role [vision] in omp now, let's see what happens. :-)
- surgical_fire 1mo agoI am using GLM on Pi without any issues. You just create an API key. Started recently though, mostly been using GLM 5.2 for planning with DeepSeek V4-flash for implementation.
- petesergeant 1mo ago[total rewrite: their subscription code is buggy. It takes a while for paid subscriptions to show up, and for upgrades to take effect. Original comment was whining about this]
- unrvl22 1mo agotheir sub is crap. use opencode go (multiple workspaces) or wait for weights to drop
- moinism 1mo agoGoogle: Here is the next iteration of our flash model series, with a discount. please use. thx. Z.ai: Here is our next iteration, neck and neck with Fable/Sol. weights releasing in two weeks.
- postatic 1mo agoLook, GLM, Kimi, Deepseek and Qwen should just join forces and come up with THE model that will beat the frontier lab models even just for the benchmaxxing perspective - all just to create hype and chaos to derail the trillion IPO conversations surrounding OpenAI and Anthropic.
- libertas_quae_s 1mo ago[dead]
- Jacopos311 1mo agoThis looks very interesting indeed!
- felixlu2026 1mo ago[dead]
- bertili 1mo agoThis will be roughly on pair with Kimi K3, but using a third of its parameters. Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in less than a month Z.ai have cut the parameter/RAM barrier to a third. Congratulation to Z.ai and all the hard working Chinese researchers who are quitely boiling the frog.
- cmrdporcupine 1mo agoCongrats def in order but as usual the proof will be in the pudding of actually running the thing. GLM 5.2 has token efficiency problems. It's not a stupid model, but it takes a lot of "thinking" to produce not-stupid results. ("But wait..."). Which makes its pricing deceptive. I tried to get by through the month of June on just GLM 5.2 and it was ... fine-ish for about two weeks. But the provider situation wasn't ideal.
- dannyw 1mo agoHave you counted your thinking tokens for say Opus or Fable? It wouldn’t surprise me if frontier closed models “over-reason” just as much, but you don’t see it thanks to the summariser. (We do know GPT5.6 have adopted the caveman shorthand, which explains its token efficiency).
- cmrdporcupine 1mo agoI have a $200 monthly Codex plan. I never run out of budget and it's... disturbingly smart. It's very hard for anything to compete with that right now. I do occasional experiments where I cancel or downgrade that and try to live on open models only and it just never works out financially or skills wise. There's nobody offering K3 etc at rates that end up being significantly cheaper. Yet.
- rammler 1mo agoKimi is a great model but it was clear from the start they achieved they brute forced that performance through scaling. The frontier models K3 compares to are rumoured to be smaller also. GLM on the other hand is way ahead in perf/parm but severly compute bound. Now once GLM can scale up or Kimi optimizes the training more, that gonna be fun times.
- ofjcihen 1mo agoThe capabilities of open models approaching or meeting that of SOTAs is good in every way except for our short-sighted economic reliance on their success (in the US at least).
- bsenftner 1mo agoSo, "cyber capabilities", whoa there horsey, what the fuck is that? Are we making up words or are you trying to court the black hat crowd?
- danggggg 1mo ago[dead]
- matheusmoreira 1mo agoMeanwhile, my OpenAI TAC application lingers in a total limbo. I suppose I'll switch to this at some point.
- alienbaby 1mo agoOne htought I had; if The chinese allow unfettered access to cyber capabilties while th US does it's best to neuter it's model releases, from China's point of view they have the US all tied up in knots dealing with problems they don't give people the tools to solve. China giggles as it watches the US under threat from people using it's models. The US is restricting citizens from owning this particular kind of weapon, while China is handing it out to the wrolds citizens freely. It feels like the US would only come out worse overall?
- onlyrealcuzzo 1mo agoI suspect Anthropic wanted the US gov to ban Mythos for marketing. If it turns out to be bad for them, the US gov will likely suddenly unban models.
- swalsh 1mo agoI suspect mythos demonstrated a fully autonomous offensive hack in a similar way Open AI's models performed, and the government is reacting to it the same way we reacted to blackhat. The threat is real.
- bigwheel 1mo ago[dead]
- smurf9852 1mo ago" a judge agent then attempts each task to verify that it is actually solvable " I understand you need to verify the goal is achievable. But if the judge agent has the same goal as the training agent (solve), and both are of the same model, then aren't the judge and the training agent doing the exact same thing? What is the point then? Can someone explain this to me.
- deleted 1mo ago[deleted]
- alightsoul 1mo agoIt works because models are not deterministic so you can prompt two instances of the same model to be adversaries of each other
- maxkim869 1mo ago[dead]
- leobuskin 1mo agoI bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent as a defender (following HF story)! I understand that such models can be used by malicious actors, but it’s fair to have it publicly available (and play on your side in case of emergency). This is what changes the world in a better way, I think, not the guardrails.
- api 1mo agoThey're invaluable for developers to fix their code. This is definitely an area where AI decisively beats human devs in a very valuable way. It can try so much surface area so fast. If it won't attack my stuff, it won't help me build my stuff to be secure.
- leobuskin 1mo agoExactly my CoT! I hope z.ai won’t change this behavior after training it on our input the same way as Anthropic did (shame on you, folks, seriously)
- jermaustin1 1mo ago> I understand that such models can be used by malicious actors, but it’s fair to have it publicly available I feel like there should be some mechanism to prove you own the code/app/site/whatever and it will remove the guardrails from the LLMs allowing them to find and fix these vulnerabilities.
- maxdo 1mo agoThey just ignore in their benchmarks opus 5 for some reason :) also grok 4.6 . I wonder why
- bigyabai 1mo agoOpus 5 has nerfed cybersecurity performance, making it hard to benchmark: https://support.claude.com/en/articles/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet https://support.claude.com/en/articles/14604842-real-time-cy...
- deleted 1mo ago[deleted]
- deleted 1mo ago[deleted]
- yogthos 1mo agoI'm so glad I managed to get their subscription when it was on sale for 250 bucks a year back when it was 5.1. Back then it was just ok, but after 5.2, it's become my main workhorse. And 5.3 is looking fantastic.
- Satoshi_Bro 1mo ago[flagged]
- jadbox 1mo agoNo API yet? I don't see it on OpenRouter yet.
- andai 1mo agoWe got nukes capable of having existential crises, before GTA 6...
- CuriouslyC 1mo agoThese results look pretty good, given the smaller model size and the GLM family's historic robustness. Cheaper than Kimi and more robust than DeepSeek. The question in my mind is if you're going cheap, are you going to stop here or go all the way down to DeepSeek Flash?
- Ruca_AI 1mo agoSame base model, this much improvement just from post-training is kind of insane. Really curious to see how GLM-5.3 performs on messy, real-world repositories once the weights are released
- jjice 1mo agoAm I correct in understanding that this is just 730B-ish parameters as an MOE? That sounds like incredible performance per parameter. The new Deepseek was also very impressive with its 280B or so. Plus the most recent 30B-ish Qwen and Muse. I find the performance to size ratio of these models to be way more interesting, selfishly because it makes me bullish on what I'll be able to run on a machine I own over the next few years. The progress is just incredible.
- zaj00l 1mo ago[dead]
- fcanesin 1mo agoGLM-5.3 is further proof that all >1T models are currently undertrained. I was looking at inteligence density ( https://www.pasteboard.co/6q2-5f92mtj9.png https://www.pasteboard.co/6q2-5f92mtj9.png ) from recent open models (where parameters sizes are known) and taking DS-v4-flash as upper limit GLM-5.x can 3x its performance.
- jameshart 1mo agoSo is ‘cyber’ just short for ‘cybersecurity’/‘cyberwarfare’ now? That is not what cyber used to mean… This is like when ‘crypto’ started meaning cryptocurrency.
- jrflo 1mo agoThat's how language works, it's always evolving...
- deleted 1mo ago[deleted]
- jamesponddotco 1mo agoIs there a plan somewhere that gives access to Kimi K3 and GLM-5.3? I was thinking of testing both to run security reviews of my code. I know OpenCode Go has both, but their limits seem kinda low, so I'm not sure how feasible it is to run such a task with them.
- scottfits 1mo agowhat i appreciate most about this post is the level of transparency in how they built and scaled an RL pipeline. my friends at the big labs are so cagey about everything, and Zai is just putting out a great crash course for free.
- ikari_pl 1mo agoSuch a smart model and didn't warn them how confusing the headline is to anyone who understands what "cyber" means?
- lazarus01 1mo agoI’m using deepseek v4 flash to build a complex full stack production ai app and it’s a total beast. I break out Claude when I hit some serious roadblocks, but that doesn’t seem to be happening much after the last deepseek flash release. Deepseek prices just went up, but are still low. I will def try GLM on my next project
- himata4113 1mo agoThere goes the last argument that anthropic had. I think beyond this point we're entering the 'dark scary world' that dario predicted which in fact result in things going on as usual. Really, the amount of fear mongering is astonishing. Hopefully they will drop it all together and focus on making models that are useful for everyone like their original mission was instead of playing games with politics.
- steffi_oliver 1mo ago[dead]
- deleted 1mo ago[deleted]
- yazan94 1mo agoHow do the new Deepseek, Qwen, and Kimi models compare against GLM 5.3?
- jsLavaGoat 1mo agoDoes it watermark?
- webbrain 1mo agofable 5 is now open-source! really!
- knguyen0105 1mo agoIn the past week, due to recent 'K3 moment', I tried both K3 and GLM 5.2. GLM worked better for me, faster and less verbose. Looking at the new improvements in 5.3, I can't wait to test it - this could be a real alternative to cc for me.
- qianyiyu 1mo ago[dead]
- ByteWarden 1mo agoglm-5.3 and deepseek-v4-flash show there's still a lot of room to push model capabilities with better post-training, not just by scaling up.
- sfn42 1mo agoWhat is "emergent cyber capabilities" supposed to mean?
- kaeluka 1mo agoIt's a smart model, but it has huge troubles staying on task... it constantly builds things I didn't ask it to. Might have its uses in a tightly controlled harness, but for coding I'm switching back to a mix of K3/GLM-5.2/Deepseek V4 Flash 0731
- kaeluka 1mo agoi came back to say: i've given this another chance and while my point still stands: it's actually making up for that by being a good thinking partner. the latest project i've been working on this has been really good. So I may have judged too quickly.
- AutomationGoat 1mo agoYeah Sonnet 5 feels drastically better.