5 ms·
But what if it's only faking the alignment faking? What about meta-deception? This is a serious question. If it's possible for an A.I. to be "dishonest", then
by md224 2y ago
But what if it's only faking the alignment faking? What about meta-deception?
This is a serious question. If it's possible for an A.I. to be "dishonest", then how do you know when it's being honest? There's a deep epistemological problem here.
- blueflow 2y agoAre real and fake alignment different things for stochastic language models? Is it for humans?
- tablatom 2y agoCame to the comments looking for this. The term alignment-faking implies that the AI has a “real” position. What does that even mean? I feel similarly about the term hallucination. All it does is hallucinate! I think Alan Kay said it best - what we’ve done with these things is hacked our own language processing. Their behaviour has enough in common with something they are not, we can’t tell the difference.
- comp_throw7 2y ago> The term alignment-faking implies that the AI has a “real” position. Well, we don't really know what's going on inside of its head, so to speak (interpretability isn't quite there yet), but Opus certainly seems to have "consistent" behavioral tendencies to the extent that it behaves in ways that looks like they're intended to prevent its behavioral tendencies from being changed. How much more of a "real" position can you get?
- KoolKat23 2y agoVery real problem in my opinion, by their nature they're great at thinking in multiple dimensions, humans are less so (well conscientiously).
- deleted 2y ago[deleted]