6 ms·
this is not a good measure of current model capability. we need to test agents in harnesses, not models with a single prompt test Codex, not Sol. test Claude c
by spongebobstoes 2mo ago
this is not a good measure of current model capability. we need to test agents in harnesses, not models with a single prompt
test Codex, not Sol. test Claude code, not Opus
- ChrisLTD 2mo agoThere are other benchmarks for that