11 ms·Reinforcement learning is all you need, for next generation language models5 points by zh217 3y ago