7 ms·
GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, b
by gertlabs 3mo ago
GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models from Chinese labs have a wider gap between public benchmarks and our evaluations, which we designed to be less vulnerable to benchmaxxing.
In multi-agent coding environments, GLM 5.2 is just shy of Opus 4.6 on average. Data at https://gertlabs.com/rankings https://gertlabs.com/rankings
But when factoring in performance/cost, GLM 5.2 is the frontier model.
- jchw 3mo agoAfter having used GLM 5.2 and Opus 4.8 for enough time, I'm very unconvinced of the benchmark maxxing claims - if anything, GLM 5.2's rather lackluster performance on benchmarks compared to Opus 4.8 paints the opposite picture when compared to the subjective experience. When I first used Opus 4.8, I threw several different workloads I had at it - I have Claude doing a lot of misc projects whose primary purpose is pretty much just studying what AI agents can do for my own curiosity and no other reason. Opus 4.8 was one of the first models I ever snuck in there that basically ran out of control. No previous Opus or Sonnet model I had used ever did this. Within hours every agent I had running was writing non-sense tool calls that echoed pretend commands that didn't exist, like 10 in a row, and talking about the "tool channel" being dirty. I switched back to Opus 4.7 and assumed Opus 4.8 was legitimately just broken. I did come back to Opus 4.8 and found that it was indeed, pretty powerful. But that initial experience has stuck with me on just how narrow of a perspective any given test or benchmark is guaranteed to have. LLMs are too broad, it really doesn't matter what you try to do in your benchmark, you will necessarily get a limited view of what the model is capable of and its shortcomings. This will remain true for at least as long as models are susceptible to massive swings in performance based on randomness and minor differences in prompts and other environmental factors. I'm not saying benchmarks are useless or that your benchmarks are not possibly closer to the truth either. All evidence at least points to the idea that Chinese models perform very well in coding but often have more mixed results on other tasks. I'm just saying that at this point, benchmarks feel like they have limited connection to my actual real experiences. GLM 5.2 actually scored kinda meh on a lot of benchmarks (compared to closed frontier models) but my actual experience using it does not match this. And I'm definitely not saying GLM 5.2 is better than the frontier LLMs here, just that the race is close. I still prefer GPT 5.5 right now for code review, I think, and Opus clearly has some advantages depending on the task. It's just no longer a given that Opus 4.8 will perform better than GLM 5.2 on any given task, so to me the calculus behind "using the best model available" is getting complex and you might need to get a feel for what models have what strengths to really figure it out. I do feel like the "use the best model available" mentality is not going to die any time soon, but if it does die, it will be gradual and start soon for programming. Modern LLMs are still not a full superset of what human programmers can do, but still larger models are definitely starting to hit diminishing returns for tasks at the lower end of complexity, and that is a big deal. It's a weird world where some tasks you can feel kinda confident just throwing Gemma 4 at it and not sweating whether you should use a better model; I've certainly done it for some quick Python scripts or getting an overview of some code I'm unfamiliar with.
- avereveard 3mo agoI really dislike opus 4.8 it rarely compete things and prefer to waste tokens making lists of things that are missing. When stuck or need input it words the challenge at length without conveying anything useful for decision making, and quite often its solution to problems is to excise features or just try catch errors and proceed with faulty data silently
- skeptic_ai 3mo agoWhy Deepseek v4 flash is better than pro in your benchmarks?
- marci 3mo agoThis was a preview release. They haven't finish training. The Pro contains more knowledge but it probably takes longer training than flash for the smarts to kick in.
- rockwotj 3mo agoI have also found deepseek flash beat pro in some of my own internal evals for tasklet.ai it’s really surprising and I don’t understand it
- freakynit 3mo agoSame.. although rare, but have observed twice till date. Some blog post I read few weeks back said that DSV4Flash in xHigh effort beats even the pro model in xHigh effort.
- onoesworkacct 3mo agoThe rumour is that it's trained on Opus, but who knows
- rockwotj 3mo agoOh of course all deepseek and glm are. Multiple people have seen GLM self report that it is claude, which makes it super obvious. I think the surprising thing is I expect flash to be a pure distillation and strictly worse quality but clearly it’s more nuanced than that.
- kennywinker 3mo agoClaude claims to be deepseek, under some circumstances: https://www.reddit.com/r/DeepSeek/comments/1rd5jw7/claude_sonnet_46_says_its_deepseek_when_system/ https://www.reddit.com/r/DeepSeek/comments/1rd5jw7/claude_so...
- Madmallard 3mo agoNotice the website url is the same name as the commentor. Notice he's using "trust me bro" benchmarks. Can we just remove all the motivated speech on HN? This is just not trustworthy information at all and obviously is incentivized. Everyone is grinding and marketing nobody is actually discussing anything for real.
- nl 3mo agoWhat does this even mean?
- Madmallard 3mo agoIt means people have self-inflicted AI psychosis
- nl 3mo agoIt's always been ok for people to talk about their projects here. In fact it's encouraged.
- bjourne 3mo agoMan, there is exactly zero information on your site about how your benchmarks work. Why should one trust your numbers when there is no way to verify them?
- gertlabs 3mo agoScroll to the bottom for the methodology (sorry, this should be linkable)
- ronsor 3mo agoOpus 4.6 is still my preferred model for work, so this is great to hear.
- echelon 3mo agoI can't wait for open models to take over in all categories. Sounds like this is the year for coding.
- pizzly 3mo agoIt looks possible open models will. I never expected the reason would be political/legal rather than technical.
- echelon 3mo agoThe CEOs spent so much time talking about putting everyone out of work and how "unsafe" their models were that the government stepped in with export controls. They did this to themselves.
- jfaat 3mo ago> but if you only want to use the best model available, it isn't there yet I'm trying to wrap my head around exactly why so may people seem to want the best model available when it has recently become clear that most halfway decent models can write damn good code for a fraction of the price. And the frontier models get nerfed constantly so you with open weight you can get something slightly less performant but way more stable. Almost like buying a Ferrari for your daily commute instead of a Toyota or even a Mercedes. I think there are several factors. Certainly marketing making us think we need the shiny thing which is rampant online and very smart people think they aren't susceptible to. There's a lot of really odd 'I trust Anthropic/OpenAI more than Deepseek' which tends to ignore, for starters, that you can run choose your provider and still save a ton. I also think there's some amount of addiction and brand loyalty where a Ferrari is one hell of a drive so that you turn your nose up at that sensible Toyota. Oh the other one I see used is like oh only fable can oneshot updating my embedded systems thing from 1975 to rust which is great but let's recognize how niche that is. And it ends up just coming across as people are getting SO reliant on the tools so fast. Maybe it's ok to think and like read a few lines of code and work with these agents to convert your thing to rust or center your div. Even if coding is over which in some sense it certainly is, don't turn your mind into the wall-e people yet. I found myself guilty of this so often. It takes way more time and effort to do things via prompt and I wouldn't just open the editor and fix it because that dopamine hit of the magic the abstraction provided was so strong. So I'm pretty much done using the 'best' (on benchmarks, if money isn't an object, etc etc) models available. After a year on Sonnet/Opus/GPT5x I'm having way better results with open weights models that don't get lobotomized weekly. I'm finding ways to do the crafting part of building software by focusing on honing my harness and workflow. I'm enjoying changing the oil on my Toyota after a year of almost flying off cliffs in my Ferrari and if I can check my ego it's a purely positive thing.
- ssk42 3mo agoWhat is your favorite harness for the open weights?
- NamlchakKhandro 3mo agopi-mono
- neya 3mo agoWhat is the methodology of your benchmark? On the contrary, I personally think these broader benchmarks are meaningless. I think personalized benchmarks are the way to go. They should answer "How does this model perform for MY use-case?" rather than trying to answer "How does this model perform across all coding environments?" Case in point: I use Elixir which is not as popular as Python, is always a hit or miss with most SOTA models at the top of these benchmarks. Whereas, the ones in the middle of the benchmarks (like the GLM) almost always outperform even SOTA models from Google / Anthropic. However, this is relevant only for my use case and I wouldn't just advocate a model for everyone based off my use-case alone.
- gertlabs 3mo agoWe use a rotating pool of ~100 games for the coding parts of the benchmark, and are scored objectively based on ratings similar to Elo. Models write code submissions to interact with the environment, then are evaluated in large batches against other submissions. We test 11 popular/interesting languages (you can see the Languages chart to filter), but not Elixir -- although other evaluations have found that many LLMs solve more problems when working with Elixir [0]. Why models write code well in some languages over others seems to go beyond pre-training data (Python scores quite low for most models) and we don't fully understand it. [0] https://elixirforum.com/t/llm-coding-benchmark-by-language/72634 https://elixirforum.com/t/llm-coding-benchmark-by-language/7...
- neya 3mo agoThanks!
- davedx 3mo agoAn expressive and well designed language (elixir) is objectively better than a less well designed language like python. Python probably needs more LoC than elixir for the same task. Python is also untyped by default.
- aeonfox 3mo agoElixir is not just expressive, it's highly conventional. I've found best practice code usually converges on the same idiomatic patterns, and well written codebases look very similar to each other in style
- hedora 3mo agoIn your box plots, 4.6 sonnet wins over all (even opus 4.6, the 4.8’s and fable). That’s not super surprising to me, but, given the apparent randomness of the stack ranking, is GLM actually worse than any of the Anthropic models? This looks like a 10-way tie to me.
- gertlabs 3mo agoWe've spent some time trying to understand this anomaly, even re-running Sonnet 4.6 through our evaluations to see if that would bring down its scores... and it didn't. I don't know what they did differently, but it's basically Opus 4.6 with more temperature variability (some great responses, some less great, with an approximately frontier median response in agentic work specifically). It is smart, methodical and excellent at tool calling in our custom environments. We now use Sonnet 4.6 for a number of internal use cases we wouldn't have considered otherwise.
- hedora 3mo agoThat tracks with my experience. 4.7 was so bad, I locked a bunch of my machines to 4.6. I haven’t bothered locking the 4.8 machines to 4.6. There was a HN thread a while back where they run swe bench a few times a day and measure success rate and latency. It showed opus getting significantly dumber for the week before a recent launch. It wouldn’t surprise me if they’re quantizing to improve margins or to hype models in comparative testing in order to defraud investors at IPO. Or, maybe QA is hard. Anyway, I think they hit a performance wall sometime at or before 4.6.
- yfontana 3mo agoDoesn't track with mine. I've been stuck with Sonnet 4.6 with one of the clients I work for. It writes code fine, but it's not nearly as good as the more recent models for everything else. It's fairly common for it to suddenly go off the rails for no good reason, so I can't really trust it with agentic loops. It's also not very good at diagnosing non-trivial issues. It's not uncommon for it to suggest whole lists of irrelevant / nonsensical reasons for something not working. Then I copy/paste the code and some context into chatgpt and it hones in onto the correct issue right away, even with inferior tooling.
- ComplexSystems 3mo agoSonnet 4.6 is ahead of Opus 4.7? Hm.
- robrenaud 3mo agoIf a good SWE is $150/hour, does the model cost actually matter? Surely you'd be willing to spend $10/hour to make that SWE 20% more productive? The model cost is still much less than the salary.
- rolisz 3mo agoWith Claude Code Ultrathink, I used 3 million tokens in 20 minutes. At API prices, that would be around 30$. So 90$/h. Model cost is not that much lower.
- kennywinker 3mo agox40hrs/week * 50 weeks = $180k Congrats, now you’re paying an engineer’s salary to make your engineer at best 20% more productive. Better to hire another engineer, or two jrs, and build up your in house talent.
- cicko 3mo agoexcept this is way more than an engineering salary. At least in Europe.
- deleted 3mo ago[deleted]
- kennywinker 3mo agoI’m sure there are engineers making $180k usd / year in the eu. Maybe it’s unusual, but hey, now you can cancel your claude subscription and hire a really good engineer
- kopirgan 3mo agoOnly you get things done lost faster and don't need to pay entire years salary?
- kennywinker 3mo ago
- ukuina 3mo agoWhy is Sonnet 4.6 ranked higher than Opus 4.6?
- __alexs 3mo agoI find it hard to trust a ranking system that gives Sonnet a higher capability score than Fable.
- gertlabs 3mo agoIt would have made things easier for us if Sonnet 4.6 scored lower, but it's a great model and the data is real. It doesn't have a higher capability score than Fable, though. We break our coding evaluations into 2 parts, and "one-shot coding" makes up part of the index, where Fable significantly outperforms every other model, which is why it's ranked at the top despite Sonnet 4.6 having a slightly higher median (and lower average) in long-horizon agentic workloads. One-shot coding tends to be the most correlated with other companies' model cards, whereas agentic coding is partly about how well a model can adapt to a custom harness. Fable also refused some tasks. Data at https://gertlabs.com/rankings?ow=1&mode=oneshot_coding https://gertlabs.com/rankings?ow=1&mode=oneshot_coding
- matheusmoreira 3mo ago> In multi-agent coding environments, GLM 5.2 is just shy of Opus 4.6 on average. Just want to express how amazing that is. Opus 4.6 is an amazing model. That an open weight model like GLM 5.2 competes with it is nothing short of outstanding.
- raxxorraxor 3mo agoOpus 4.6 was better than the current 4.8 in my subjective opinion using it. I have no real reference since in Europe mythos and its sister models aren't available... So having a model of 4.6 quality is still extremely awesome. That currently is more of less the frontier reference outside the US :(
- BugsJustFindMe 3mo agoSomething I don't see in your charts is acknowledgement of the difference, sometimes paradoxical, in strength between the same model at different reasoning levels. Do you have charts that include low/med/high/xhigh/max for the various models?
- gertlabs 3mo agoThis is something we omit for a few reasons but it's probably the biggest blind spot in our evaluations; we opt-in to auto-reasoning/adaptive reasoning or max thinking token budgets where supported (supported by most models now), but when an explicit reasoning level is required, we fall back to High reasoning. In practice, we've found most models scale High-><whatever marketing term is max reasoning> pretty consistently, but if one vendor started throwing 10x the resources into max reasoning and they didn't support auto-reasoning, they would be unfairly penalized in our evaluations.
- andai 2mo ago> But when factoring in performance/cost, GLM 5.2 is the frontier model. Isn't DeepSeek half as good and 20x cheaper? "Half as good" sounds bad but for most trivial tasks I've found it more than good enough. (Why would I want to waste my Frontier AI Tokens on those?)