5 ms·
You can prove that doing this will spiral training into a fixed point. There was a lot of research into getting this to work in the past, but it never truly wor
by hodgehog11 11d ago
You can prove that doing this will spiral training into a fixed point. There was a lot of research into getting this to work in the past, but it never truly worked well. The hope was that if RLVR was used quite a bit, and the general performance crossed some threshold, that it would then be possible. However, since it has been shown that RLVR only concentrates the distribution of outputs rather than truly shift it, I doubt this will ever be a viable strategy.