6 ms·
TL;DR is that they didn't clean the repo (.git/ folder), model just reward hacked its way to look up future commits with fixes. Credit goes to everyone in this
by sabareesh 9mo ago
TL;DR is that they didn't clean the repo (.git/ folder), model just reward hacked its way to look up future commits with fixes. Credit goes to everyone in this thread for solving this: https://xcancel.com/xeophon/status/2006969664346501589 https://xcancel.com/xeophon/status/2006969664346501589
(given that IQuestLab published their SWE-Bench Verified trajectory data, I want to be charitable and assume genuine oversight rather than "benchmaxxing", probably an easy to miss thing if you are new to benchmarking)
https://www.reddit.com/r/LocalLLaMA/comments/1q1ura1/iquestlabiquestcoderv1_swebench_score_is/ https://www.reddit.com/r/LocalLLaMA/comments/1q1ura1/iquestl...
- ofirpress 9mo agoAs John says in that thread, we've fixed this issue in SWE-bench: https://xcancel.com/jyangballin/status/2006987724637757670 https://xcancel.com/jyangballin/status/2006987724637757670 If you run SWE-bench evals, just make sure to use the most up-to-date code from our repo and the updated docker images
- LiamPowell 9mo ago> I want to be charitable and assume genuine oversight rather than "benchmaxxing", probably an easy to miss thing if you are new to benchmarking I don't doubt that it's an oversight, it does however say something about the researchers when they didn't look at a single output where they would have immediately caught this.
- domoritz 9mo agoSo many data probes would be solved if everyone looked at a few outputs instead of only metrics.
- alyxya 9mo agoGiven the decrease in the benchmark score from the correction, I don't think you can assume they didn't check a single output. Clearly the model is still very capable and the model cheating its results didn't affect most of the benchmark.
- stefan_ 9mo agoNever escaping the hype vendor allegations at SWEbench are they.