6 ms·
https://paperswithcode.com/sota/code-generation-on-humaneval https://paperswithcode.com/sota/code-generation-on-humaneval
by PodgieTar 3y ago
https://paperswithcode.com/sota/code-generation-on-humaneval https://paperswithcode.com/sota/code-generation-on-humaneval
- riku_iki 3y agoHuman Eval is very different to SWE-Bench on which Devin is tested
- PodgieTar 3y agoI didn't say it was the same, I compared non-agentic Claude to this. This used HumanEval.
- riku_iki 2y agoYou said: > how this performs against the same benchmark Devin was using > ... > Claude 3 Opus already scored around 85-86% on these benchmarks Devin used SWE-bench, not HumanEval, which kinda implies you said Opus got 85% on SWE-bench which is not true. This was my confusion..