9 ms·
OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than t
by theptip 12d ago
OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt.
Anthropic have also observed similar things, so while it seems to me that OpenAI’s level of control is more of a dumpster fire, it’s by no means a unique issue to them.
- Sharlin 12d agoIt's almost as if it's not actually possible to align an unknowable mystery box of floats.
- hansvm 11d agoGood thing we're not trying to deploy them into fully autonomous weapons or anything....
- theptip 11d agoI wouldn’t take the fatalistic stance that it’s fully impossible - but it’s certainly impossible to align a model while racing as fast as any technological paradigm shift has ever raced.