7 ms·
Speculative decoding does this to an extent - using a smaller model to generate its own predictions and putting them in the batch of the bigger model until they
by aiddun 2y ago
Speculative decoding does this to an extent - using a smaller model to generate its own predictions and putting them in the batch of the bigger model until they diverge
https://huggingface.co/blog/whisper-speculative-decoding https://huggingface.co/blog/whisper-speculative-decoding
- brrrrrm 2y agoIt doesn’t. It simply trades compute efficiency by transposing matrix multiplications into “the future.” It doesn’t actually save FLOPs (uses more) and doesn’t work at large batch size
- imtringued 2y ago>doesn’t actually save FLOPs (uses more) Does anyone even care? Really, who cares? The truth is nobody cares. Saving FLOPs does nothing if you have to load the entire model anyway. Going from two flops per parameter to 0.5 or whatever might sound cool on paper but you're loading those parameters anyway and gained nothing.
- brrrrrm 2y agocompanies that run these things care - they run at huge batch size and are compute bound in the limit