6 ms·
do you even needs thumbs up? I've been long suspecting that code models get better because they use our data and our results from feedback, status codes, green
by lackoftactics 7d ago
do you even needs thumbs up? I've been long suspecting that code models get better because they use our data and our results from feedback, status codes, green tests for reinforcement learning
- user43928 7d agoI was thinking that the training is more curated, so that the methods are learned from experts, and that measurably successful behavior is reinforced. Throwing in random chats with some sentiment analysis doesn't seem like the most promising method to me, but I can only speculate.
- lackoftactics 7d agohmm, maybe not sentiments, running commands can produce binary results to reinforce, but that's also speculation