5 ms·
Anthropic specifically calls out that example as >"illegible reasoning in a few reinforcement-learning environments over long rollout" Yet, I get the point th
by ewjt 9d ago
Anthropic specifically calls out that example as
>"illegible reasoning in a few reinforcement-learning environments over long rollout"
Yet, I get the point that you're making: those tokens essentially are an internal scratchpad for the LLM which isn't required to logically lead to the output.
This video presentation of the paper you linked was interesting:
https://www.youtube.com/watch?v=hUp3zh23aHw https://www.youtube.com/watch?v=hUp3zh23aHw