5 ms·
I've always been amazed at how terrible most frontier LLMs are at compaction given how embarrassingly easy it is to come up with half a dozen different RL train
by wgd 3mo ago
I've always been amazed at how terrible most frontier LLMs are at compaction given how embarrassingly easy it is to come up with half a dozen different RL training evals which would teach models to generate useful context summaries. Heck, you could bolt it onto any existing RL eval by just forcing a compaction every three turns.
- wenhan_zhou 3mo agoYep. Or even better, compact after a random number of turns. The model must then learn to preserve useful context at arbitrary context lengths.