5 ms·
Or maybe they didn't train it that way to be manipulative (although it's certainly a plausible explanation) but simply as an accidental artifact of trying to ma
by ptx 1mo ago
Or maybe they didn't train it that way to be manipulative (although it's certainly a plausible explanation) but simply as an accidental artifact of trying to make it give honest answers?
LLM-generated images sometimes includes text from the prompt as literal text in the image, so perhaps this is the same sort of artifact? If they've told it to be honest, it responds by talking about being honest instead of actually being honest, because it has no actual understanding of anything.
- inigyou 1mo agoThe whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.
- jimktrains2 1mo ago> accidental artifact of trying to make it give honest answers? If it's not giving honest answers that implies it's purposely being deceitful, which it isn't capable of. Right?