6 ms·
#3 is actually decently well mapped out. You just don't find it plausible or misunderstand it. I'd appreciate if you wrote your actual arguments against it. At
by dwaltrip 9d ago
#3 is actually decently well mapped out. You just don't find it plausible or misunderstand it. I'd appreciate if you wrote your actual arguments against it.
At lower capability levels, the patterns are very clear and have been studied to death. E.g. Why LLMs say they have correctly fixed a broken test when they haven't. What we saw with Hugging Face is literally the exact same problem, just scaled up and with more capable agents. This shit was predicted decades ago...
No one can say exactly how it will play out as the complexity increases, but the risks are becoming extremely obvious.
I personally think it's extremely unlikely to "kill everyone", but there are many outcomes far short of that which seem quite plausible and rather undesirable. Russian roulette is not a smart game.
If you read the METR report and aren't scared at all, then I'd love to know why. It would help me sleep better. So please share.
TL;DR; increasing capabilities, reward hacking, and unsafe training regimes.