8 ms·
Some critical issues with the SWE-bench dataset
- acc_297 2y ago> 32.67% of the successful patches involve cheating as the solutions were directly provided in the issue report or the comments. Is this what Hofstadter means by a strange-loop?
- andrepd 2y agoTurns out "AI deep research reasoning agent" was just "we can print the training set"
- thegeomaster 2y ago...by piping it through the world's most inefficient echo function.
- sva_ 2y agoThat reminds me of someone calling the Bitcoin blockchain the most expensive linked list in the world.
- wongarsu 2y agoThe difference is that Bitcoin is designed to be "just" an append-only* timestamped linked list, with some rules on how a new node can look like in order to be successfully appended. Making the creation of a canonical linked list possible between hostile actors is the whole innovation. The currency stuff is "just" a cool practical application tacked on the linked list LLMs by contrast are not designed to just repeat what's already in the instructions, no matter which stance on LLM design you subscribe to * exceptions apply
- xrd 2y agoYou should immediately publish a paper on Arvix with your revolutionary IEF brand, an improvement on transformers and mamba architectures. Then, like Ilya, take $1B in funding the following week.
- modeless 2y ago> When we filtered out these problematic issues, the resolution rate of SWE-Agent+GPT-4 dropped from 12.47% to 3.97%. This matches my intuition about the coding performance of these models a lot better. I don't think any current coding benchmark accurately measures coding performance.
- alfalfasprout 2y agoYep anecdotally that's basically spot-on. It's also one of the reasons that I still find copilot vastly more useful than highly autonomous AI tooling (cursor, roocode, avante, etc.)
- OsrsNeedsf2P 2y agoAnecdotal but I was always shocked to see Claude 3.5 perform so poorly in the benchmarks, when it generates 80% of my code in Cursor (and in cases it fails, no other model succeeds)
- TheDong 2y agoDifferent people seem to get wildly different results here, and I'm not sure what percentage is down to the type of software being built vs the usage patterns. In my case, I would guess less than 10% of the code I get out of AIs is useful. What sort of code are you getting those results with? Is it yet-another-react-frontend-button? Is it ebpf programs? Is it a parser in rust? For the latter two, I've found AI to have pretty low rates, and for the former I haven't had the desire to try.
- aprilthird2021 2y agoMy gut tells me the AIs will be best for small web projects that are greenfield. The kind a 1-3 person team could maintain. And my gut tells me they are the worst for the kinds of long-established software conglomerates many professionals work at, which have tons of internal services, integrated acquisitions, etc. etc. Ultimately the AI is good at what the average developer online is good at, probably full-stack web dev of projects from scratch.
- semi-extrinsic 2y agoSo what we need is something like a versioned crowdsourced coding LLM eval dataset. Every quarter, you have a couple thousand volunteers provide 2 GitHub issues from the past 3 months, which are nontrivial to resolve, and where there exists strong test cases. Each volunteer then cross-checks 2 issues from other volunteers. The volunteers get 1 month free subscription to some AI service in return. This dataset is then published as SWE-UberBench-2025-02 or something. People can then only evaluate their coding LLM on datasets published after their training period.
- SR2Z 2y agoRight, so that AI companies can freely throw this significantly more valuable training data into a model and then turn around and advocate for clamping down on the freedom of models.
- delusional 2y agoAnd why would these "couple of thousand volunteers" help with this?
- rsynnott 2y agoAnd how would you ensure that all of them were really volunteers and not colluding with the vendors? Like, tech companies cheating on benchmarks is an old, old story (personal favourite: in the dark ages, before 3D acceleration, some graphics card drivers, on detecting a 2D acceleration benchmark, would _simply draw the wrong thing_), and I wouldn’t trust at least three of the major players as far as I could throw them.
- delusional 2y agoI'm pretty sure my bios still contains an option to "improve performance of 3dmark 8" or something similar.
- nitwit005 2y agoIf you know some way to get people to volunteer millions of dollars of free labor, there are better uses of their time than evaluating LLMs.
- otterley 2y agoI am shocked—shocked—when a vendor cheats in order to increase their benchmark scores. I always tell my customers to ignore benchmarks and compare outcomes with their own workloads. Benchmarks are almost completely useless in the real world.
- Snuggly73 2y agoI only trust benchmarks that I’ve faked myself :)
- commandlinefan 2y agoAlthough I believe there's a lot of this going on, in this case it just appears to be incompetence rather than malice.
- adamc 2y agoI don't know why you are getting downrated. That is sane advice.
- optimalsolver 2y agoYou need benchmarks with the following three properties: 1) No known solutions, so there's no "ground truth" dataset to train on 2) Presumably hard to solve 3) But easy to verify a solution if one is provided. This, of course, is easier done on the STEM side of things, but how do you automatically test creativity, or philosophical aptitude?
- hsuduebc2 2y agoI guess it's purely subjective. Maybe some internal commission if it comes to quality of creative work?
- brap 2y agoMy own impression with SoTA models is that they’re very useful for coding, yet they suck ass for solving unique problems (which is the case for every sufficiently large codebase).
- ukFxqnLa2sBSBf6 2y agoThere’s a few things I’m not understanding here. 1. Did the benchmark authors not review the issues and make sure the solution was not present in the issue? 2. Are the issues locked after they’re included in the dataset? You’d think they would be immutable for reproducibility. 3. For the agents writing patches, is test running part of their inner loop validation? If they write a patch that makes the test pass, then the jobs done. Or is that validation step kept secret from the agent? I don’t see how unless the tests aren’t part of the repo.
- jbellis 2y agoEspecially with swe-verified, I thought that was the whole point of that dataset
- flakiness 2y agoThis was also my first thought, but reading [1] again, what they did was labeling like: > Whether we consider the issue description to be underspecified and hence unfair to be testing on. > Whether the FAIL_TO_PASS unit tests filter out valid solution and a bit more. This is pointed out in the linked paper too. The moral of the story to me is that, don't believe the paid human annotator. You can (hopefully) still believe the PhD students doing these unpaid jobs as their research ;-) [1] https://openai.com/index/introducing-swe-bench-verified/ https://openai.com/index/introducing-swe-bench-verified/
- sebzim4500 2y ago>1. Did the benchmark authors not review the issues and make sure the solution was not present in the issue? I looked at a bunch of issues in the dataset when SWE-verified first game out and I was trying to make scaffolding to solve it and I don't remember a single time where the solution existed verbatim in the issue. I'm not saying it never happens, but it would have to be rare. > 2. Are the issues locked after they’re included in the dataset? No one changes the issues in the dataset but of course the original issue on github will have been resolved long ago. The models don't have access to this in their context, but if they were trained on github there's a very real risk that they've seen the solution. > 3. For the agents writing patches, is test running part of their inner loop validation? If they write a patch that makes the test pass, then the jobs done. Or is that validation step kept secret from the agent? I don’t see how unless the tests aren’t part of the repo. The tests aren't provided to the model, they are run after the model has proposed its final answer.
- MattDaEskimo 2y agoThere's a serious issue with benchmarks. Instead of resolving it, some leaders are further complicating their meaning Such as OpenAI grading their benchmarks based on "how much money they made" or "how easy a model was convinced to hand over fake money".
- huac 2y ago> 32.67% of the successful patches involve cheating as the solutions were directly provided in the issue report or the comments. Looking at the benchmark, https://www.swebench.com/ https://www.swebench.com/, about half of scored submissions score under 1/3 correct? So they're either not cheating, or not cheating effectively?
- nraynaud 2y agoyeah, in the abstract they demoted the score from 12% to 3%, so sadly retirement is not yet here :(
- sebzim4500 2y agoLLMs do not reliably reproduce their training data. This is quite easy to demonstrate, every LLM has been trained on all of wikipedia (at minimum) and yet there if you ask it a niche fact mentioned once on wikipedia it is highly likely to get it wrong.
- huac 2y agothat comment refers to the test time inference, i.e. what the model is prompted with, not to what it is trained on. this is, of course, also a tricky problem (esp over long context, needle in a haystack), but it should be much easier than memorization. anyways, another interpretation is that the model needs to also make a decision on if the code in the issue is a reliable fix or not too
- sebzim4500 2y agoThen I don't understand what he's suggesting. It is obviously not the case that 1/3 of the questions int he SWE-bench dataset have the solution in as part of the issue that is provided to the model. You can just download it and look. The solution is likely in the training data though.
- feznyng 2y agoThis is why I’m a bit skeptical of the o3 results. If it’s spending a bunch of time reasoning aren’t the chances of it simply regurgitating a solution it saw in its training data at some point in its output stream higher? It still needs to be clever enough to identify it as the correct answer but it’s not as impressive as an original solution.
- bearjaws 2y agoI would argue almost every popular benchmark quoted by the big LLM companies is tainted. OAI, xAI, Antropic, Google all score incredibly well, then you go to try and write code and its just okay. They claim it can do PHD level reasoning, but here I am not trusting it on basic computational thinking.
- jandrese 2y agoYeah, that's true in many fields with these AI agents. They demo well, but when you put them to actual work they fall right on their face. Even worse, the harder the task you set for them the more they lie to you. It's like hiring a junior dev from one of those highly regimented societies where it's more important to save face than to get the job done.
- aprilthird2021 2y agoYour last sentence feels kind of spot on. The lack of transparency around confidence in the answer makes it hard to use (and I know it would not be simple to add such a thing)
- hackernewds 2y agosounds like a skill issue to be honest. you could probably tell the assistant to just ask you questions when information is missing instead
- ryoshu 2y agoProgramming is easy. Asking the right question is hard. People don't know what questions to ask.
- aprilthird2021 2y agoBut it doesn't know when information is missing
- dimitri-vs 2y ago
- htrp 2y agoPaper from October 2024
- shayanh 2y agoI found that this paper was submitted to ICLR, but got rejected: https://openreview.net/forum?id=pwIGnH2LHJ https://openreview.net/forum?id=pwIGnH2LHJ To me the analysis of SWE-Bench is a solid contribution and informative. My guess is that to meet conference's submission bar they had to come up with their own bench (SWE-Bench+), which wasn't thorough enough and the paper got rejected mainly because of that.
- vonneumannstan 2y agoAcceptance or rejection at big ML Conferences doesn't seem to carry much signal either way anymore. Completely saturated by grift and poor quality so each paper should be evaluated independent of their Conference status imo.
- OldGreenYodaGPT 2y ago> solutions were directly provided in the issue report or the comments This is fine, many of my real tickets already explain the solution. A good ticket often offers a solution or where to start looking.
- softwaredoug 2y agoYep that's fine for an issue, but a problem if you're trying to eval whether AIs can solve coding problems.
- ionwake 2y agoI was wondering how long this would take to surface, you can tell a surprising amount just by carefully watching how the trainers answer interview questions, which is kinda meta really.
- comex 2y agoSome of the examples in the paper seem to be wrong. For django-31056, they claim the AI-generated patch is "incomplete" because it's "missing critical parts of this logic, such as the try-except block and the check for a running event loop.". But if you look at the diff, that's clearly wrong. The try-except block and running check were already there before the patch. The human patch just indented them, making them appear as both - and +, while the AI patch didn't. To me, the AI patch seems correct. It's slightly less efficient than the human patch when DJANGO_ALLOW_ASYNC_UNSAFE is set, but slightly more efficient when it isn't (which is the common case!). The human patch does feel more natural, but the AI patch is fine. I'd grade it a tie between human and AI. For django-32517, they claim that the human and AI patches "produce entirely different outputs", but actually they do exactly the same thing. The human version has `reversed(self.dict)`, while the AI version has `reversed(self.dict.keys())`. `reversed` treats the object as an iterator, and iterating over a dictionary in Python just gives you the keys, so it doesn't matter whether you call `.keys()` first. The human patch is more idiomatic, but it's also more confusing, as shown by the fact that it confused the authors of this paper. I'd grade it another tie. Edit: I tried to sign up for OpenReview so I could leave a comment about this, but the system wouldn't let me register without completing a form that assumes you have an academic position. Perhaps I should email the authors.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- fourpostmaun2 2y agoThe entire premise of this paper is false. They claim that the "hints_text" is used and leaks the answer in Section 2.1.1; however, the authors of SWE-Bench themselves state that this is not used anywhere (Issue #133 on the official SWE-Bench GitHub). According to the paper: > 1. Solution leak: represents instances where the solution to the issue is clearly outlined in the issue description or comments on GitHub. Since both the issue descriptions and comments (referred to as hints_text in the SWE-Bench study) are provided as input to the models, these LLM models can extract the solutions directly from this information instead of generating it independently. And yet, the SWE-Bench authors themselves explicitly state: > In short, for participating on the SWE-bench leaderboard, using hints_text in any manner is not allowed. Although we don't explicitly say this in the original paper, we also do not make any mention of using the hints_text anywhere. So, it's a made up issue that would only occur if you deviated from the paper implementation and explicitly added a field called "hints" that isn't used anywhere.
- perrygeo 2y agoThe solution moving forward has to be private benchmark suites. I could see teams investing in their own set of programming challenges and periodically re-evaluating them - similar to how we would construct sets of live interview questions for candidates and qualitatively assess their ability. It's so vital that it's not leaked and that it's fit-for-purpose and manually assessed. These general purpose, public benchmarks based on questionable metrics are effectively worthless to assess real programming skill. Case in point, as others have mentioned here, Claude scores modestly on these benchmarks but vastly better than the alternatives in practice. I don't trust Claude fully but far more than OpenAI models; it's not even close. The IRL performance advantage is not reflected in any of these benchmarks.
- 1024core 2y agoTo quote Goodhart's Law: When a measure becomes a target, it ceases to be a good measure. Or, as in the case of LLMs and benchmarks: When a benchmark becomes a target, it ceases to be a good benchmark.
- deleted 2y ago[deleted]
- dang 2y agoSubmitted title was "SWE-Bench tainted by answer leakage; real pass rates significantly lower". Normally we'd replace that with the article title, in keeping with the site guideline ("Please use the original title, unless it is misleading or linkbait; don't editorialize."), but in this case the article title is so generic that this is arguably misleading as well, so I took a representative phrase from the abstract instead. That's preferable, because it's better to use the authors' own representation of their article. If anyone can find a better title (i.e. more accurate and neutral, preferably using language from the article itself) we can change it again. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- alalv 2y agoSomething weird (or at least uncommon) that has caught my attention and I havent seen mentioned in the comments is that they cite the swe-bench paper author by first name in the abstract, Carlos et al, and then by last name (as it is usually done) in the paper, Jimenez et al.