6 ms·
Prefilling any model with the first 1% of reasoning tokens from another model should always move the output towards the output of the other model directionally
by Fripplebubby 1mo ago
Prefilling any model with the first 1% of reasoning tokens from another model should always move the output towards the output of the other model directionally - that's just next token prediction doing its thing.