6 ms·
Because distillation-only is a quick performance shortcut that only lets you get to the level of the thing your distilling for the most part or a little bit wor
by novok 12d ago
Because distillation-only is a quick performance shortcut that only lets you get to the level of the thing your distilling for the most part or a little bit worse and does not allow you to actually progress past it. It's like only being able to make VHS copies of videos, and maybe do some basic video editing without being able to actually go out with cameras and make new movies.
To actually have something competitive and improved within the next 3 months and not be perpetually behind, you need your own independent model creation process. So to extend the metaphor, a complete movie studio with cameras, actors, staff, sets, budgets, etc. It's the right strategic move to do when you are GPU constrained, which the Chinese labs are, but it won't let you get past it.
A bunch of pedantic people will come out of the wood work citing a bunch of things saying that is not the case because of some detailed mechanics of how model training works and they will get fixated on some of the words I used, but zoom out to the level of what an AI lab is able to produce and this becomes evident.
- nvme0n1p1 12d agoThat doesn't really answer the question. What they do internally for training the next model is a separate issue. I'm talking about the models they offer publicly. Per the article, companies are dropping OpenAI+Anthropic (partly) because of costs. If distilling is so simple and easy, why doesn't OpenAI take this "quick shortcut" and serve a self-distilled model externally, so they can charge reasonable prices and stop bleeding customers? Surely they can at least match the Chinese labs' efficiency, right? Wouldn't more customers and less opex look good for the IPO?
- linkregister 12d agoAnthropic, OpenAI, GDM, and Meta spend more on training than other labs by an order of magnitude. If they felt safe reducing this spend they would. These labs fear getting outcompeted.
- nvme0n1p1 12d agoAgain, I am talking about inference, not training. Please read.
- linkregister 12d agoHow do you think they fund training? This is just as asinine as insisting that drug manufacturers only price medications based on production costs.
- nvme0n1p1 12d agoLess opex = more profits. They can use that money to fund training. I don't see how that could possibly be a bad thing.
- linkregister 12d agoI indeed failed to understand your point. Isn't that what they already do with Claude Haiku, GPT-5.6-Terra, etc?
- nvme0n1p1 12d agoThat's the goal, but those smaller models don't match the price:performance of leading open-weight models, which is why companies are switching away (as explained in TFA). The open weight labs figured out some secret sauce that (so far) big name labs are unable to replicate, so instead of competing, they're going on the defensive with claims of distillation attacks.
- linkregister 8d agoAgain, in the friendly article, only one company is on record. The article allows the reader to make whatever inferences she may want to make.