8 ms·
I have also found deepseek flash beat pro in some of my own internal evals for tasklet.ai it’s really surprising and I don’t understand it
by rockwotj 3mo ago
I have also found deepseek flash beat pro in some of my own internal evals for tasklet.ai it’s really surprising and I don’t understand it
- freakynit 3mo agoSame.. although rare, but have observed twice till date. Some blog post I read few weeks back said that DSV4Flash in xHigh effort beats even the pro model in xHigh effort.
- onoesworkacct 3mo agoThe rumour is that it's trained on Opus, but who knows
- rockwotj 3mo agoOh of course all deepseek and glm are. Multiple people have seen GLM self report that it is claude, which makes it super obvious. I think the surprising thing is I expect flash to be a pure distillation and strictly worse quality but clearly it’s more nuanced than that.
- kennywinker 3mo agoClaude claims to be deepseek, under some circumstances: https://www.reddit.com/r/DeepSeek/comments/1rd5jw7/claude_sonnet_46_says_its_deepseek_when_system/ https://www.reddit.com/r/DeepSeek/comments/1rd5jw7/claude_so...
- trentor 3mo agoDon't ask western llms in Chinese what model they are...
- xbmcuser 3mo agomaybe they distilled claude for the flash version and not for the other hence better tool use and programming benchmarks