9 ms·
Seems just slightly better than last v4 release, considerably (3x) more expensive, but also faster and slightly more token efficient. https://aibenchy.com/comp
by XCSme 7d ago
Seems just slightly better than last v4 release, considerably (3x) more expensive, but also faster and slightly more token efficient.
https://aibenchy.com/compare/deepseek-deepseek-v4-1-flash-high/deepseek-deepseek-v4-flash-0731-high https://aibenchy.com/compare/deepseek-deepseek-v4-1-flash-hi...
- gunalx 7d agoIgnoring the obvious ai slop webpage. I don't really trust the benchmark. It seems either pretty saturated, or inconsistent just based on the results.
- XCSme 7d agoWhat seems inconsistent? The coverage is quite small, only 22 tests. It's more to compare the cost/speed/consistency between models, given the same tasks.
- gunalx 6d agoRight. I got the feeling of it being saturated because all the top 5 fully completed it.
- XCSme 6d agoYeah, it's hard to find a single simple task that all models fail on, in low context length conditions. Also because models now are actually not that good on knowing things (domain knowledge), as they rely more on web search on tool use. So if I added a question, about some obscure fact, probably the SOTA models would fail it, but in practice they would find it with web search enabled. Not sure how to handle that. This is also why Gemini is on top, it's good enough at coding and instructions following, while having by far best general and domain specific knowledge.
- sinuhe69 7d agoYour benchmark is a curious one. I didn't see you included Muse Spark 1.3 contributor even though its price is much lower even than DeepSeek. The low price changes many recommendations completely. And FWIW, DeepSeek retain and train on your data, too.
- XCSme 7d agoI waited until NovitaAI provider became available on OpenRouter. I have their guardrails enabled to not allow requests to providers that train on data. As far as they say though...