5 ms·
For those who want to dive deeper, here’s a 300 LOC implementation of GRPO in pure NumPy: https://github.com/superlinear-ai/microGRPO https://github.com/superli
by lsorber 1y ago
For those who want to dive deeper, here’s a 300 LOC implementation of GRPO in pure NumPy: https://github.com/superlinear-ai/microGRPO https://github.com/superlinear-ai/microGRPO
The implementation learns to play Battleship in about 2000 steps, pretty neat!