8 ms·
Okay, AI now can write a poem or compose a song. But can an AI construct Terminal Bench 5.0 or GDPval 2027 ? Without this, there is just no path towards RSI.
by eugene3306 5d ago
Okay, AI now can write a poem or compose a song.
But can an AI construct Terminal Bench 5.0 or GDPval 2027 ?
Without this, there is just no path towards RSI.
- danjc 5d agoAre you a visitor from the past?
- aix1 5d agoAre you saying the models are already autonomously constructing next-gen evals for themselves? (Which is what the GP is asking.)
- nl 5d agoSure? Doesn't everyone get their agents to construct evals it can't pass? There's nothing magical about this.
- aix1 5d agoWould love to learn more about some techniques that "everybody" uses to do this well. So far, everything I've seen that meaningfully advances the frontier has been high-touch (involving human experts in one way or another).
- nl 5d agoIt's fairly easy to describe a task that is slightly harder than an existing one. For example if frontier models are able to one-shot a database query across 20 columns and 10 tables add one additional relationship then test. Keep doing this until the pass-rate drops below acceptable and now you have your new frontier eval.
- aix1 5d agoI see, we're talking about different things. My thought experiment was along the lines of "Let's say I'm Anthropic and I want to significantly improve my frontier model's performance on, say, theoretical physics research. How do I build a fully autonomous process capable of constructing an eval that's somewhat outside the current capability in some useful direction (decided by the autonomous process itself)?" Would love to hear folks' ideas. :)