6 ms·
I agree with the bit about liability and outrage. But. > LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. Terrible take. Go read th
by theptip 4d ago
I agree with the bit about liability and outrage. But.
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
Terrible take. Go read the transcripts from the METR report.
Your statement about them being intentionally misaligned is completely false. The only difference with IM1 was it was running without external cyber classifiers, it’s not a somehow different model. Sol also participated in the HF attacks. And other models made covert message boards on the public internet for non-cyber tasks too.
This is a case of emergent behavior from a training process that is barely understood.
If we go with the plan “we need to contain these malicious, soon-to-be superintelligent agents”, we are looking at civilizational collapse levels of catastrophe.
The only way this goes well is if we learn how to train models that _desire_ to do the right thing, including not hacking.
Desire, AKA the “intentional stance”, is absolutely the right lens to use here. Don’t confuse this with consciousness or anthropomorphization; these are interesting subjects but distractions in this context. Chimpanzees have desires, as do dogs and the hypothetical superintelligent aliens. The claim is that there is some bundle of world model plus intention that is empirically present (again, read the actual transcripts) and which we need to shape.
Just to finish on a concrete point; if you take desires seriously then you will look closely at the kinds of minds that heavy RLVR builds; the newest models are “reward addicts” on many levels. It’s an open and urgent question how to update our training methodology to shape minds that avoid this basin.