5 ms·
whats about self-consistency like in grpo with majority voiting?
by oedemis 1mo ago
whats about self-consistency like in grpo with majority voiting?
- kumama 1mo agocastform founder here. i'm personally a little against techniques like self-consistency/majority voting during rl training because they tend to result in the model's output distribution "sharpening" a lot. this means the model will lose it's exploration ability and probably won't be able to explore/discover new solution strategies, which can be harmful for both rl training + generalization to unseen cases