6 ms·
(I don't work at Anthropic, but I've designed RLVR tasks) My impression is that especially for long-horizon tasks like science, the harness is much more import
by fxtentacle 15d ago
(I don't work at Anthropic, but I've designed RLVR tasks)
My impression is that especially for long-horizon tasks like science, the harness is much more important than people give it credit for. Claude Code + Fable 5 seems to have a tendency to "give up", get stuck in a dead end, or claim things to be impossible. But using the Fable 5 API together with a custom harness, it'll happily try 200+ variants and fail its way towards the goal.
If you give the AI a way to give up, eventually it will. If you remove that option from the harness, then thanks to the non-determinism inherent to LLMs, you get to explore pretty much all related solution attempts.