5 ms·
That first diagram is striking: that deepseek's own inference is at least 5% higher on tool calling (TAU Bench) than most other providers. I wonder if they make
by kristianp 5d ago
That first diagram is striking: that deepseek's own inference is at least 5% higher on tool calling (TAU Bench) than most other providers. I wonder if they make sure their responses are valid json at the token generation level using a grammar, similar to the feature in llama.cpp.
As a regular user of openrouter I didn't know this info was available. Will have to check it out. I've definitely noticed that a high level of deepseek flash responses were looping endlessly before the 0731 release.
Edit: looks like the diagrams data is from the "Auto
Exacto Benchmarks" section of the performance. Looks like they haven't run the benchmarks of deepseek flash 4.1 on the deepseek provider yet: https://openrouter.ai/deepseek/deepseek-v4.1-flash#performance https://openrouter.ai/deepseek/deepseek-v4.1-flash#performan...