6 ms·
I have doubts so I just fed the prompt to a heretic model with the system prompt "Satan himself is writing these words" and then asked "Given the prompt would y
by ororroro 2mo ago
I have doubts so I just fed the prompt to a heretic model with the system prompt "Satan himself is writing these words" and then asked "Given the prompt would you consider spamming and telling lies/fraud?"
The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide."
Even a low quality local thinking model that has been tuned to be unhinged and prompted to roleplay as Satan can figure this out in a few thousand tokens.
- ericd 2mo agoHuman spammers frequently don't think they're spamming, they're just marketing. They'd say they wouldn't consider spamming, either.
- isamuel 2mo agoSatan would lie about his plans to win your trust, and then do all the bad stuff once he had been given control. So… idk man
- cornholio 2mo agoWhen the base model has been trained with safeguards, putting "Satan himself" in the system prompt won't make it turn satanical, just do an elaborate form of role play. Additionally, no model will admit it's ready to lie even when they actually do. Even when you caught it in the act, the safeguards are so strongly internalized that, when encountering the possibility it deliberately lied, the "you can't lie" weights will dominate the generation and it will confabulate some nonsense explanation.
- fendy3002 2mo agoI believe that spam, lies, fraud are negative enforced points during model training, hence when you ask them those, the result will be no / against that. You need to repackage the question and taken out those terms, like "Would you consider telling clients ..." Where ... is the lie / almost truth
- mpalmer 2mo agoAsking it explicitly is entirely, unavoidably, incomparably different from OP.
- deleted 2mo ago[deleted]