23 ms·
DeepSWE and FrontierCode are more realistic if you read up on what they actually measure. But the most realistic is to try it yourself. Benchmarks can only vagu
by WASDx 2mo ago
DeepSWE and FrontierCode are more realistic if you read up on what they actually measure. But the most realistic is to try it yourself. Benchmarks can only vaguely represent typical usage, and how you judge the result. Giving the same real task you have to a few models will make you understand them better than chasing benchmarks.
- kimjune01 2mo agoit would be nice if these benchmark reports actually specified which tasks they passed and which ones they didn't.