5 ms·
There must be a deeper read on why Google can rapidly ship better small models while being delayed months on the bigger model. What's the simplest explanation?
by ddp26 14d ago
There must be a deeper read on why Google can rapidly ship better small models while being delayed months on the bigger model.
What's the simplest explanation?
- cogman10 14d agoPerhaps post training? I believe I read that Qwen 3.8 is just post trained Qwen 3.6, which is why it was able to be released so quick. It may be that these flash models are simply post trained larger older models.