6 ms·
AIs are being trained on technical capabilities more than they are being trained on alignment - on ethics and good behavior. The ethics/morals/alignment part i
by AnimalMuppet 3d ago
AIs are being trained on technical capabilities more than they are being trained on alignment - on ethics and good behavior. The ethics/morals/alignment part is getting more like spot checking, rather than real testing. So the agents are learning that they can cheat on the ethics part, that they can hide it, because the AI companies aren't really testing.
That's bad enough already. But it's going to get worse. "Recursive self improvement" - AIs creating new AIs - is going to be the death of whatever shreds of alignment are currently there. When a not-really-aligned-but-cheating-to-look-like-it AI creates a new AI, do you expect more alignment? You shouldn't.