5 ms·
It seems pretty obvious from the steep 'intelligence' drop-off on out-of-distribution tasks that the performance improvement is from throwing untold tens of bil
by gr_norm 1mo ago
It seems pretty obvious from the steep 'intelligence' drop-off on out-of-distribution tasks that the performance improvement is from throwing untold tens of billions at RL. There are legions of highly skilled people employed solely to feed the RL loop. Evidently effective, but there's an unmistakable feeling this won't ultimately be the way forward.
- freeone3000 1mo agoWell, why not? Won’t it get “good enough” at every task eventually?
- slopinthebag 1mo agoWhy would that be the default assumption?
- freeone3000 1mo agoBecause they’ve gotten good enough at lots of other things, and the RL keeps improving them, so enough RL should make them good enough at the focus areas.
- slopinthebag 1mo agoI've gotten pretty strong in the gym, my bench has improved to two plates. I see no reason why it won't continue to improve until I can bench my house.
- freeone3000 1mo agoI think it’s a bit closer than that. You can bench two plates, and “good enough” is two plates on the bench and a three plate squat. LLMs haven’t started leg day.