10 ms·
Its remarkable how Anthropic is able to maintain their edge against all competition. Anyone have any idea what the secret sauce is that has Anthropic at the top
by czhu12 2mo ago
Its remarkable how Anthropic is able to maintain their edge against all competition. Anyone have any idea what the secret sauce is that has Anthropic at the top of all leaderboards for the past few years?
- x312 2mo agoGiven their pricing, I'd guess their models are just way bigger in parameter count. They've always underperformed in cost-per-performance. They also target a cost-insensitive market (corporate/coding users) compared to Google/OpenAI which support massive amounts of free users.
- hn1986 2mo agobecause in the real-world, it's far better than the rest. That's why few people use Grok, it's not even close in day to day work.
- Handy-Man 2mo agoFrom what I have read, their pre-training team is much better than anyone else. For OpenAI, their post-training team is better. And apparently OpenAI has consistently struggled at training a bigger model than GPT 4 level
- sulam 2mo agoI’m a VP Eng — the backend team I manage strongly prefers CC and Opus. The Android team I manage strongly prefers Codex and GPT 5. I’m personally not sure that the answer doesn’t just come down to stylistic differences in prompting and ergonomics in the harness. The folks that prefer Codex seem to get better one-shot results, whereas those that prefer CC are doing more iterative prompting. At any rate, I don’t think you should write OpenAI off when it comes to coding.
- jeffybefffy519 2mo agoIts even different than that, some Codex models like 5.3-codex are terrible at front end work but excel at backend/system design.
- hello_newman 2mo agoI think it's focus? Anthropic seemed to double down early on being more business/prosumer focused. While OAI, Gemini, Grok, etc were also doing various side quests like image generation, Anthropic seemed to only focus on 1 thing, and that seemed to pay off
- nullbio 2mo agoSomeone has to know. Would be nice if an insider would drop some hints so that the open-source space could make some good progress.
- ben_w 2mo agoNobody has to actually know the secret of their own success, especially not relative success to equally-secretive near-peers. Same as with rich person autobiographies: even when they tell you what they think it is, they can't see the path not travelled.
- nullbio 2mo agoI'm purely talking about the technology - not their business strategy. I actually think their business strategy is blatantly obvious and atrocious.
- steve1977 2mo ago> Same as with rich person autobiographies: even when they tell you what they think it is, they can't see the path not travelled. Yup, there's a lot of survivorship bias in those. And humans want to attribute success to some skill somehow. You cannot just have been lucky.
- nijave 2mo agoMy gut feel is Anthropic is very technical and pedantic which makes their models really technical and pedantic. They're top at code and technical benchmarks but anecdotally I've found OpenAI to be significantly farther ahead for general usage. Opus 4.8 will burn 10k tokens trying to answer something 100% whereas GPT-5.5 will burn 2k getting it 90% which is good enough for many things. Some personal testing on a "help me find that restaurant" prompt https://gist.github.com/nijave/2873b8b10d8c732e46264237b075507a https://gist.github.com/nijave/2873b8b10d8c732e46264237b0755...
- enraged_camel 2mo agoThe problem is that the remaining 10% can bite you in bad ways. I was in Cotswolds, UK a couple of months ago. For those of you who don't know, it's a rural region known for its "chocolate-box" villages and honey-colored limestone architecture. Basically, you go from village to village, most commonly via bus, taking in the sights and doing touristy stuff. When planning the trip, my sister used ChatGPT, which helpfully (and relatively quickly) found the bus schedules and times for each hop. Midway through the day, though, we ran into a huge problem: it turns out bus schedules are different on Sundays, and more limited. Which meant we couldn't actually go to our primary destination (the Model Village), and had to cut the trip short. Yes, ChatGPT was quick and pleasant to use, but missed a crucial detail. Afterwards I tried it with Opus and it did not make the same mistake.
- nijave 2mo agoArguably I'd call that the 90%. In my case, answering the restaurant question correctly with "Rishi" in my tests was the sole intent and 90% of the problem. All the models "helpfully" added extra junk about the closure, dates, quotes, etc and many of them got these details wrong--the 10% or extra crap not central to the question. If the central question was "what is the bus schedule on `day`" and the model screws that up, it gets a fail in my book. Also curious if Google Maps gets the timetables correct (assuming it has them). Semi-related, I also discovered that the default web search/fetch tools are pretty primitive and Exa MCP annihilates them. I ended up doing some comparisons with Claude Code comparing built-in server-side to Exa and to a Python MCP that used SearXNG for search and Exa was a clear winner and Python+SearXNG ended up coming out roughly the same after a few cycles of letting Claude optimize the Python code and adjust SearXNG settings. Ultimately it landed on this (making some changes to optimize returning relevant context directly in the search results so the model didn't need an additional web fetch call) https://gist.github.com/nijave/604c43e3e0fdcd60f5280d3a6b1096e9 https://gist.github.com/nijave/604c43e3e0fdcd60f5280d3a6b109...
- small_model 2mo agoI think it's the talent, laser focus on single product set and being early so ahead, same with Open AI who are only a sliver behind. Google, XAI are the next level down but they have other concerns.
- bredren 2mo agoI think it is a mix of the sibling replies here. I'd add that the company has seemed to find ways to ~do more with less. I have never liked the various nerfs Anthropic has used to balance GPU (slowing down responses, quota variance, model optimizations etc) and it definitely has burned a lot of good-will. But it has seemed that being able to look beyond the short term pitchforks has worked quite well.
- hectdev 2mo agoI think they have a better agent personality which pushes back and isn't sycophantic. It has been awhile since I've used the others but that's where it locked me in and I've stuck with it.
- giancarlostoro 2mo ago> isn't sycophantic Not sure about that one... But I think the true secret sauce for all these models is how they reason. GPT never outputs how it thinks, which "saves on tokens" but Claude absolutely tells you how it thinks, and there's people who use how it reasons about solving problems to finetune smaller open source models, with surprisingly better output.
- hectdev 2mo agoFrom my experience, it has not been sycophantic in the sense that it pushes back and questions my own reasoning in healthy ways. There were moments where I felt I was brushing up against actual AI psychosis, and it pushed back on my questioning of its intentions, that it even had intentions. I'll put it this way: I feel comfortable recommending Claude to people who haven't experienced AI yet. As we've learned from early experiences with other models, leading people down paths of believing they understood math in ways nobody else has and even harming themselves, I put Claude as a safer alternative.
- hraxz 2mo agoI think Opus can still have sycophantic residue that Fable can point out sometimes. Both models though hold their ground so well. I have got so use to the Claude personality / style of conversation that I really can't be bothered to try these other models anymore. They need to take a huge jump but that seems to be getting harder and harder because of the jumps Anthropic makes. This Grok version is a joke if it is not even clearing the bar now. I am just getting use to and using Fable more and more. I am also trying not to forget that this is the highly delayed old Fable model that Grok can't even beat on release. There will be a new version that expands the lead in a week or two. It all harder and harder to judge too. I just had a prompt/response this morning that Fable finally displayed its intelligence and vowed me. That is partly because anything with even the vaguest reference to biology defaults back to Opus.
- levocardia 2mo agoI think the "secret sauce" is not juicing the benchmarks. Claude models just feel like they are better than the benchmarks suggest, in terms of smarts and creativity, while models from every other company feel worse relative to what you'd think from the benchmarks. Only company to really internalize Goodhart's Law, IMO.
- solenoid0937 2mo agoYeah every model has great benchmarks. Claude is the only model I want to use when I'm not worried about the marginal cost of tokens (which is most of the time at work.) I then use cheaper models like GLM for personal projects but they're noticeably much worse despite being similar in benchmarks.
- deleted 2mo ago[deleted]
- mnicky 2mo agoOne angle could be their interpretability research? They understand what's going on in LLMs probably much better than anyone else. This must somehow pay off. I think it's not only an alignment/security tool but could perhaps be used for capabilities as well.
- logicchains 2mo ago>Its remarkable how Anthropic is able to maintain their edge against all competition. Anyone have any idea what the secret sauce is that has Anthropic at the top of all leaderboards for the past few years? It's self-reinforcing: they've got the best coding/research model, which helps them to improve their models better than the competition so they stay ahead.