42 ms·
Because they are an integral part of the solution generation. By distilling, you are training a model via Reinforcement Learning (RL) to mimic the answer of a
by Otterly99 1mo ago
Because they are an integral part of the solution generation.
By distilling, you are training a model via Reinforcement Learning (RL) to mimic the answer of a bigger model. To do that, you need all the steps that contributed to generating an answer.
To give you a better idea, imagine teaching a student how to solve math problem:
1. You give it the problem and the answer only (no thinking trace)
2. You give it the problem, the intermediary steps and the answer (full transcript)
I think you can agree that the second method is more likely to give a well-informed student.
- maleldil 1mo agoIsn't distillation done through SFT by having the student model approximate the teacher's output distribution?