7 ms·
I'm supposed to believe that DeepSeek distilled a 300B-1T param model with 150,000 requests? Lol.
by colingauvin 23d ago
I'm supposed to believe that DeepSeek distilled a 300B-1T param model with 150,000 requests? Lol.
- brookst 23d agoAre you thinking distillation goes from zero to complete model? I believe it can be used in the RL / fine tuning sense, in which case 150,000 requests, assuming every one was detected, could move the needle in quality. I agree it couldn’t replace all of pre and post training , but I don’t think that’s the claim. You do typical training, then distill really difficult cases.