6 ms·
There's some cool research that looks at how strongly the weights are aligned through training vs adherence to the system prompt. Like when you know a model is
by throwaway219450 11d ago
There's some cool research that looks at how strongly the weights are aligned through training vs adherence to the system prompt. Like when you know a model is lying through censorship: https://arxiv.org/html/2603.05494v2 https://arxiv.org/html/2603.05494v2
Presumably if negative guidance is in the system prompt, there's a good chance that the model would happily comply if it wasn't there.