9 ms·
Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
- Uptrenda 4d agoreal software engineering benchmark is how much stress you can take at work. Everyone knows this, bakka.
- dgellow 4d agoA bit of a meta question: what are the most relevant benchmarks by now?
- andriy_koval 4d agoNvidia and OpenAI claimed AGI, but you still have a job.
- tetec1 4d agoEpoch.ai has a global score and tracks many benchmarks: https://epoch.ai/benchmarks https://epoch.ai/benchmarks
- redox99 4d agoTerminal bench 4 is good largely because it's recent so it hasn't been benchmaxxed yet. It's more of a sysadmin/devops benchmark than a coding benchmark though, but still a decent proxy. https://artificialanalysis.ai/evaluations/terminalbench-v4-0 https://artificialanalysis.ai/evaluations/terminalbench-v4-0
- taintech 3d agoI think openrca https://github.com/microsoft/OpenRCA https://github.com/microsoft/OpenRCA is quite interesting
- jcmontx 4d agoI’ve been able to offload most tasks (coding or eles) to Codex since 5.3-codex with extra high thinking
- riddlemethat 4d agoAstra lets me offload entire projects without worrying about individual tasks…
- jeffybefffy519 4d agoDo you review the outputs?
- what 4d agoCan you show us some of these of projects?
- hattimaTim 3d agoCan you share us your workflow? I am getting dumb outputs even with Astra max.
- traceroute66 4d agoSo TL;DR benchmarking in a completely non-reproducible manner ? "Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company". So basically pinky-promise benchmarking ? I'm not sure I follow the value here ?
- kadoban 4d agoIf it builds up history and perceived reliability, this type of thing can be valuable. You're giving up transparency for it being harder to game.
- traceroute66 4d ago> You're giving up transparency for it being harder to game But then if we take that argument to its natural extreme, surely it means people should take the marketing bullshit published in the 100-page system cards published by Anthropic & co as "valuable" too ?
- kadoban 4d agoI think you know that's basically nothing like this? The model cards have every incentive to be biased, this doesn't necessarily. But even so, pretty much yes: companies that actually have reliable and accurate info in their releases get trusted more. It takes time because the default is to disbelieve info from biased sources, but it is possible to trust some of them more than others.
- cbg0 3d agoAs long as the ones offering the benchmark aren't trying to sell you something and have no affiliation with one of the companies on the page I'll take it as opposed to having the benchmark rendered useless in 3 months when the next models drop.
- demibabs 4d agoDoesn’t it ultimately have to be this way, to prevent saturation?
- bix6 4d agoWake me up when September ends or when I can do this locally.
- demibabs 4d ago> Each task comes from a private production codebase that we licensed from a real-world company How does that work?
- gruez 4d agoFrom the same site: https://withspecific.com/company-data https://withspecific.com/company-data
- traceroute66 4d ago> How does that work? My gut feeling is that any serious real-world company with a proprietary codebase worth looking at would not be handing out the crown jewels to a third party. License or not. I don't doubt somebody licensed their codebase to them, I just have my doubts about who the "who" could be.
- InsideOutSanta 4d agoCode isn't worth all that much if you don't own the associated IP, mainly copyright. And even if you disagree with that premise, if you trust that they can keep the code secret, it's basically free money. At any rate, I'm not sure it matters whose codebase it is. I'd even say that a shitty codebase might make for a better test.
- traceroute66 3d agoThe point I'm making is that companies who are serious enough to want to keep their code-base in-house and off the various online repo services are also the kind of companies who are strict about what you can and cannot do with LLMs (if they permit use of LLMs at all). So it does not make sense that the same companies would then magically sign-off on allowing their entire codebase to be spoon-fed into a whole bunch of LLMs for benchmarking.
- strobe 4d agolot of ads everywhere offering to buy your codebase of real product/star up even it long gone or failed (offer usually price per lines of code). So most likely that they have bunch of abandoned codebases between small and medium sizes and probably also some fake codebases as well.
- ahmetaytar 4d ago[dead]
- IshKebab 4d agoI think these benchmarks are not that useful, e.g. this suggests Fable is better than Astra, but in practice Astra is waaaaaay faster (like 5x; it's not even close), and also waaaay less annoying to talk to. There's only two or three sane options here - you can easily try them all and pick yourself.
- rovr138 4d agoThey're not measuring speed nor annoyance. It's there on the page
- coderenegade 4d agoI switched from Claude to Codex because Claude just doesn't do what you actually tell it to half the time. It dances around the edges and does busy work without actually tackling a tough problem. I'm not sure what others are doing that they're getting such different results, but I'll take Codex every day of the week.
- janaksunil 3d agothat's fair - for long horizon engineering tasks would speed still matter?
- IshKebab 3d agoI'd say so. Do you want your results in a week or a day?
- bdlowery 4d agoThe fact that gemini 3.8 flash is so high up there just tells you this is an awful benchmark. Try and use gemini 3.8 yourself for any real world work and you'll see it's terrible. It'll just go in circles reading the same file 20 times for no reason making hundreds of tool calls for a simple change. EDIT: I was using gemini cli... it's not a harness issue lol
- thereitgoes456 4d agoWhy so brazenly confident? Isn’t it possible that the benchmark is correct, and your experience is correct too, but you haven’t tried all the thousand different modalities of work that programming encompasses and so maybe you don’t actually have standing to judge?
- 0x457 4d agovery outdated experience from me: when I first tried gemini something, in an existing rust codebase, it looked around for files that would indicate if its go, javascript, java or c++ project, then declared I must have asked it build a new app in javascript and proceeded to circle around to figure out how it can install node and npm on my machine. So I totally believe that Gemini is just bad. Which is surprising because Gemma is very good for some tasks, but I never ever had any success with Gemini, be it in cli or chat thing or anything else that has gemini branding.
- deleted 4d ago[deleted]
- siddbudd 4d agohavent tried that model, but it sounds like a potential harness issue. Have you tried it in different harnesses?
- tucnak 4d agoHard disagree. I use 3.8 flash in Antigravity a lot, and thoroughly prefer it to most Pro-class models. It's really fast, and I've had it make crazy progress on compiler-like problems that previous models including Opus simply failed at. On ultra plan you can have it going for hours, and make incremental progress with good prompting for review interrupts. It solved a problem I couldn't solve for weeks in under 6 hours. 10k LOC total. The harness and test suite is key.
- lmeyerov 4d agoMy intuition is that many of the better & bigger 'private' code bases, at least in terms of claude code and codex... are not in fact private at this point. One lesson of running botsbench.com, in a slightly different domain, is to measure for model contamination every time.
- kwamenum86 4d agoI spent a while in big tech and remember several unique patterns of internal code based. Your comment inspired me to try to coax ChatGPT into spitting out code that was inspired by proprietary, private code. Surprisingly, it did it with no problem - I referenced an idiom from a tech company and it wrote code that really would have only been relevant for that vertical. When I asked how it learned the pattern, it said “from my learned understanding of CompanyX’s internal coding conventions”. When I asked “how do you know about those internal conventions” ChatGPT said “I don’t have access to that internal code, I overstated what I know”. Internal coding conventions are the least of our worries at this point, cat is out of the bag.
- esikich 4d agoWell how do you know which statement is truthful? These LLMs confidently say they know things that they don't all the time.
- Grimblewald 4d agomuch harder to do in OP's case, matching flavour then referencing that specific companies name when asked how it know to flavour this way? thats astronomically low for randomly selected plausible tokens without some data prior, like that companies codebase. My own experience is opus being lousy at an extremely niche math task, but it was still easier for me to describe what it needed to do to get code and correct issues in its reasoning/working than to write myself. a minor model number change later and it's nailing everything, despite my opt-out. Its is astronomically unlikley others were working on this also, especially at that level, especially this application. so, safe to say they _all_ train models on chats, the only difference being if you "opt out" you at least have some defence later when they steal your work and claim it as their models original output.
- visiondude 4d agothis is the closest benchmark to my experience using the model harness combo. Astra for as great as it is falls slightly behind Fable 5.1 for me for large feature work (although it comments code much better). in particular, Fable is able to assess priority better than Astra (meaning Astra sometimes does things that aren’t worthwhile while missing things that are clearly important, particularly on possible ballooning scenarios- fable catches “this works for x amount of data but if we run this on y way greater than x amount of data we’ll run into issues). Gemini 3.8 is under appreciated, use Google Stitch to see it in action if you haven’t used Agy yet.
- ShellfishMeme 4d agoAstra constantly does this for me. It goes 90% of the way with some task but then skips the most important part. Then when told to please fix that and do it properly, it suddenly goes down a rabbit hole for 6h and fixes scenarios that aren't even relevant. It's awful at assessing what is important to do and what not, and where to ask for permission and where not.
- majormajor 4d agoDo you find Fable significantly better than Opus at avoiding-overengineering? All of my recent testing of Anthropic models seems like they're tuned-to-hell to (a) be much slower than they need to be (running tests over and over during the loop vs at the end, say, even if those tests take a few minutes a pop) and (b) doing exactly that sort of "built a lot of fancy enterprisey feature-adjacent 'stuff'" even before nailing the actual feature. Sol and Terra both have some of the latter but they seem to do the actual work a fair bit faster (this may be a usage-based-priority-tier/rate-limit thing though) which helps offset it. I think the bigco folks saw all the "it wrote all this code but the tests didn't pass" or "it wrote the feature but it's super brittle" and tuned the newer model+harness combinations incredibly aggressively to try to turn a lazy prompt into "median Enterprise Architecture design suggestions" to bring up the baseline, but in a way that slows you down if you don't want that. I'm not on big enough subscriptions to want to burn a lot time just evaluating Fable/Astra comparatively until they're cheaper, heh. I can steer any of the cheaper ones just fine anyway.
- skilledDevelope 4d ago[dead]
- jstummbillig 4d agoI am trying to estimate if my reaction to seeing GPT-5.6 Sol last on that list is reasonable or or mostly emotional and find that I have no way of telling.
- CompoundEyes 4d agoI do think it’s the wizard not the wand at this point given a decent model. These benchmarks don’t have the wizard. Otherwise I wouldn’t see others in the exact same codebase struggle and underutilize agents while others thrive using the exact same ones.
- howunfortunate 4d agoIn other words, we're still in the era of centaur chess.
- didgeoridoo 4d agoSol failing mostly on “unverified assumptions” and rarely hitting “integration errors” seems about right to me. I think Sol is second only to Astra (and miles ahead of even Fable) in architecting & engineering the right implementation — but only if you are extremely specific and provide tight guidelines and guardrails. If you give it a one-liner… you’re going to have a bad (SHA-256-hash-verified) time.
- enraged_camel 4d ago>> I think Sol is second only to Astra (and miles ahead of even Fable) in architecting & engineering the right implementation — but only if you are extremely specific and provide tight guidelines and guardrails. To me, having to give extremely specific instructions and provide tight guidelines and guardrails defeats the purpose of agentic coding agents almost completely. At that point I might as well do the task myself. With Fable I can start with a general ask like "I'm trying to do X, can you investigate and tell me what the shape would look like" and have it poke around and think, ask me questions with single-choice or multiple-choice answers, then break the task into small chunks, each of which becomes a ticket. With Astra, it's like pulling teeth. It often does not understand what I'm trying to do, takes things literally, does not go above and beyond (i.e. infer intent), and stops way too short of the actual goal. I have to constantly prod it and it's frankly exhausting.
- prometheus1992 4d agoDoes this mean they ended up sharing those private codebases with OAI, Anthropic etc? Also, the ~30% number tracks with my experience. I thought I was going insane for expecting too much from the models but they are still bad, including astra. This morning it messed something pretty trivial while fixing an issue which I was shocked to see. Also2, benchmarks don't mean much these days.
- doctorpangloss 4d agothe requests went into a pipeline that turns them into de-identified, but salient, training data, yeah. everywhere except maybe bedrock.
- ramigb 4d agoCan you please share, if you are comfortable of course, what did the model(s) mess up? what were you using codex/cc/pi? did the project have a solid agent.md/claude.md? I am genuinely curious whenever someone have such a low success rate with models what is happening because it could be fixed maybe? From my own experience using agents for the past year or so. The rate if I have to guess, is well above 70%. I mainly use claude (opus) on typescript react projects that are well setup with minimal plugins/MCPs! happy to share more if you are interested.
- bel8 4d agoI'd love to see these: - DeepSeek V4.1 Flash - Kimi K3 - GLM 5.3 (and flash) - hy4-preview - Grok 4.6 All of these can be acessed using a $10/mo OpenCode Go subscription.
- jwolfe 4d ago3 of those are already in there.
- throwaway473825 4d agoHere's the list: 1 Fable 5.1 38.8% 2 GPT-6 Astra 33.8% 3 Gemini 3.8 Flash 31.2% 4 GLM 5.3 28.8% 5 Grok 4.6 23.8% 5 Muse Spark 1.3 23.8% 7 Kimi K3 18.8% 8 GPT-5.6 Sol 16.2% See number 4, 5 and 7.
- taintech 3d agoThank you sir!
- janaksunil 3d agowill do! happy to chat more on janak@withspecific.com as well
- obilgic 4d agoGemini 3.8 flash has been incredible for our agents. For us, It performs better than any other model except Fable.
- hollars 4d agoThe high score of Gemini 3.8 Flash vibes with my experience anecdotally. While it often goes off the rails with open-ended questions (which is a strength if taken with care), it is also a good at solving issues in a well-defined environment like an enterprise codebase.
- deleted 4d ago[deleted]
- aryansingh9034 4d ago[flagged]
- matheusmoreira 4d agoI used a similar methodology. Code review is my most requested action, so I used blind code review results to compare the frontier AIs. Even posted an article about it: https://www.matheusmoreira.com/articles/code-reviewing-lone-lisp-with-sol-and-fable https://www.matheusmoreira.com/articles/code-reviewing-lone-... Unlike TFA, the lone lisp code is public. I suppose the models could have been trained on my codebase. Still, I think it produced some interesting results. Took months and loads and loads of tokens to do this, so I'm not gonna repeat this study as new models come out. It did anchor all of my future expectations, though. OpenAI is winning as far as I'm concerned, and their cybersecurity program is the only remaining pain point.
- paidx 4d ago[flagged]
- finn888 4d agoAveraging pass@1 across eight runs per task is useful; it exposes harness consistency instead of letting one lucky resolution dominate.
- janaksunil 3d agoyes!
- andai 4d agoAGI 38.8%
- matt3210 4d agoThese'll be part of the training set eventually.
- hefu_hk 4d ago[flagged]
- freakynit 4d agoThis is the first set of benchmarks which match my observations around gemini-3.8-flash perfectly. This model is a true hidden gem.
- thefourthchime 4d agoReally? I gave it a trivial HTML job, and it went off for fifteen minutes. It did eventually did a do a decent job, but I can't wait that long.
- sureMan6 4d agoI gave it a trivial HTML job and it messed it up in several ways including being lazy and lying about results It's probably very hit or miss like everything with LLMs but I was really surprised it performed that badly
- freakynit 4d agoI'm using it through antigravity cli .. and in every single run (100's by now), this model was fast, and the outputs were of good quality. Of course, not at Astra or Fable level, but, close to like Sol-low level.
- cute_boi 4d agoThis benchmark is shitty because it puts Gemini in 3rd position. I tried gemini on simple code base and it invoked 210 tool calls just to update 3 lines of code.
- ttul 4d agoWe built a “code atlas” that provides the LLM with a semantically queryable map of how things connect and relate in a very large and sprawling codebase that evolved over 15 years. It tends to dramatically reduce the length of time models have to spend reading code while also making sure they are aware (within their context window) of nuances that are important that might be missed were they forced to just rely on reading the code in hundreds of repositories. I strongly recommend trying this approach out yourself. The recipe is not rocket science. Get your coding agent to take a first cut at building the atlas itself, and then manually correct it. Once you’re happy that it got things right, put an MCP on it or a CLI or whatever. And your LLMs will know what to do from there.
- m3kw9 4d agowhat if the production code base was made by mostly by Anthropic models?
- nullbio 4d agoGreat point.
- janaksunil 3d agoall the codebases were written pre-2023, so pre when AI got good at coding
- chandureddyvari 4d agoIt also depends on your skills and tooling (test execution and verification- agent browser, functional/unit etc) GPT 5.6 Sol lagging behind Kimi, GLM 5.3 is surprising to me. IMO Fable 5.1 ~ Astra > GPT 5.6 Sol > Opus.
- skhameneh 4d agoSome things in this seem reasonable, but others just don’t make sense and there’s crucial details missing (like reasoning levels and what harness was used). For example, I found Kimi K3 to use more tokens than some other models, which caused it to cost twice as much purely because of the token volume. This experience lines up with ArtificialAnalysis’s benchmarks, but not these. There’s a number of other comparisons here that don’t match up with my experience or other benchmarks. By many accounts, this is the outlier. I could attribute the differences to harnesses used or something like reasoning levels, but none of those details are published. While this seems interesting, I can’t take this seriously. Correction: The harnesses are listed as a column, I missed that. My other concerns and questions still remain, it’s unclear why some of their results are the outlier that does not match my experience, ArtificalAnalysis’s benchmarks, or some of the experiences of others commenting.
- arshxyz 4d agoThey seem to be using the provider's harness for each
- janaksunil 3d agowe've done our best to use the native provider's harness. all models were run on 'high' reasoning. this is still v1 and tons of room for improvement - really appreciate your feedback!
- glub 4d agoI'm surprised Sol and Astra are leading "Unverified assumption" metric and Fable is better there. I run Fable as my main model with Sol as advisor that watches every turn. Fable likes to throw around assumptions that it didn't check that are simply false, and Sol always goes to actually verify them and then alert Fable it's assuming things. I've tried reversing this pairing with Fable as advisor. It'll just sit there going "sounds good"
- ghoshbishakh 3d agoI suggest you flip them. The verifier role will always verify. You will see Sol making assumptions and Fable fixing them. But I agree - Fable makes some spectacular assumptions (which are poor assumptions).
- glub 3d agoI think what matters in this case is how proactive and greedy the model is. GPT models are extremely proactive and gredy. So when Fable mentions something that may affect some obscure component of the system, GPT will start digging the codebase, execute web searches, re-read AGENTS.md and hit fable on the head. Fable never does that, it just reads the turns and acknowledges it read them. This also explains why GPT models tend to overengineer things and why they're amazing reviewers if you triage their findings.
- manmal 3d agoGPT doesn’t do all of that all that much when it itself is the implementer. RL has made implementation and reviewing two different behavior sets.
- LewisVerstappen 3d agoWhat do you use as the harness?
- glub 3d agooh-my-pi. I was using my own homegrown (mega slop) harness for a while, but it distracted me from working on my actual projects, and I realized oh-my-pi was doing the same things I've been doing, including advisor, native server-side compaction, etc. It's a really good harness.
- arshxyz 4d agoDreadful color-coding on the output tokens table
- janaksunil 3d agowould love to learn why?
- springtimesun 3d agoI built exactly this using my own codebases. The setup isn’t that complex; split the git history to just before the change, sandbox the agent with everything they will need at that commit and lightly modify rules so they don’t go searching outside the box. Then they get the same prompt (usually the ticket that began the work) and are graded against the accepted PR. The thing that takes the most time is finding the examples. In my real dev flow it’s rarely ticket -> PR -> merge, things bounce around a lot more. So, even though the stated goal is to get away from one shots, that is basically the environment you have to set or else test for specific other outcomes (e.g. agent stopped and raised a question when it realized x). It takes time to do, but I would really recommend it. Now I can push new open models through the batteries and see how they line up to past ones in a few days (I run them locally, it’s slow). It moves my sense of x model is good at y and bad at z to from vibes to a better heuristic (these still run at temp 1, heuristic is the correct way to think about outcomes IMO). It grounds it in your actual code and problem space. My takeaway from my testing: in Rails or front end codebases, most models I test are competent and with a human in the loop they would accomplish their goal of getting to a mergeable PR. They are not as good as Claude and since I pay subsidized rates via subscription Claude still gets first pass. They are very worthwhile to layer in as reviewers and catch many issues. My anxiety about a rug pull by the frontiers has been turned way down. I would have to adapt to a local only flow, but it wouldn’t be much adaptation and the opens can deliver in their current state.
- janaksunil 3d agohey this seems really interesting - what prompted you to test multiple agents on your codebase?
- springtimesun 3d agoCuriosity and anxiety. I rebuilt my entire workflow around agents so the unease that the frontiers would change something (access, pricing, availability) and lock me out of that were high. Also why I spent way too much on hardware (at least that can be deducted). Now the whole stack could run in my house and I feel much better about the situation. Once I got the testing going though it is worth it for its own pursuit. Building processes around the dev process and trying to get the best outcomes is at least as fun to me as actually delivering client code. For the first time in my tech career I feel like I’m in a place with no maps. No one has done my experiments yet. I have a custom quant of K3 at Q5 that lets me get 10 tok/s on a CPU inference box (admittedly you need a 72GB Blackwell also). As far as I can tell no one else has done this. It’s such an exciting time!
- janaksunil 3d agoi'm janak, cofounder of Specific Labs (YC F25) and one of the authors of Real-SWE. if i can help answer any questions please feel free to email me at janak@withspecific.com, happy to send over my phone number as well :)
- option_greek 3d agoA lot of this tracks but misses the variations in what the models in general are good for.
- songhonglei1985 3d ago[flagged]
- barbegal 3d agoWithout a human to benchmark against it's really tough to gauge how good these models are vs how good the task definitions and existing codebases are. My intuition from the example full instructions are that the tasks are poorly specified which results in ~60% failures due to bad assumptions and missing requirements.
- taintech 3d agoI think currently AI models did incredible improvement against 2025 OpenRCA research paper with 11.25% success rate. Key questions, are they topped in performance? Is there some next leap?
- nottorp 3d agoThat's my impression from my job, which involves a "private, real-world"[1] codebase. You need to explicitly tell the LLM what to take into account when generating code or it will miss things. And that's with months of saved memories. Where it fails is on plugging business logic in. Tell it to generate a new piece of UI and it will do that fine. My guess is the 10x success stories are for from scratch applications where 80% of the code is boilerplate and self contained modules. And no one comes back and tells you maintenance dropped the productivity improvement to 2x. Btw since it's mentioned a couple times in the comments: the Claude Team and Enterprise subscriptions do not train on or retain your code by default. Apparently the promises were good enough for my employer and their customers. [1] By the way, real-world smells of LLM generation. Normal people write "real world".
- felixlu2026 3d ago[dead]
- quocdat25mle 3d agoVery cool work. Please also run on new DS Flash 4.1, an important oss model nowsday
- sergeyk 3d agoIf anyone wants this kind of benchmark for their own codebase, happy to set you up with superconductor.com/benchmark 1. Import your own PRs 2. We find the original spec or infer the spec 3. Agents you select (eg claude code opus 5, codex gpt 6 astra high, pi kimi k3, etc) implement the spec (starting from the parent commit of the PR, with the git history is pruned so they can’t look up the implementation) 4. Three judge LLMs grade agent solution given spec, PR implementation, and a rubric 5. You see the scores
- Betelbuddy 3d agoThe LLM vendors say their AI is going to kill us all and started calling their latest releases AGI. In the meanwhile realistic benchmarks like these ones, show they can only complete 10% to 15 % of the tasks, and now that Jon Skeet, Marc Gravell, BalusC, and Darin Dimitrov have gone on strike they will be progressively worst.