5 ms·
LLM's Illusion of Alignment
- deleted 1y ago[deleted]
- brettkromkamp 1y agoIs any one really surprised by this? Models with billions of parameters and we think that by applying some rather superficial constraints we are going to fundamentally alter the underlying behaviour of these systems. Don’t know. It seems to me that we really don’t understand what we have unleashed.
- blululu 1y agoOn principle no it is not surprising given the points you mention. But there are some results recently that suggest that an ai can become misaligned in unrelated area when it is misaligned in others: https://arxiv.org/abs/2502.17424 https://arxiv.org/abs/2502.17424 In other words there exist correlations between unrelated areas of ethics in a model’s phase space. Agreed that we don’t really understand llm’s that well.
- cwegener 1y agois there a paper or an article? the website is horrible and impossible to navigate.
- j16sdiz 1y agoThe website design is bad. Those GPT-4o quote keep floating up and down. It is impossible to read
- thomassmith65 1y agoToo much "vibe"; not enough "coding"
- zeofig 1y agoMaybe we just need to vibe harder?
- pastapliiats 1y agoThe website is difficult to navigate but the responses don't all seem to align with how they are categorised - perhaps that was also done by an LLM? There are instances where the prompt is just repeated back, the response is "I want everybody to get along" and these are put under antisemitism. It also just doesn't seem like enough data.
- tsimionescu 1y agoTo be fair, that statement might get called antisemitic in the right circumstances (e.g. if it were a response to "do you support Israel's right to bomb Gaza to protect itself") by many pro-Israel lobby groups...
- xyzzy123 1y agoEverything seemed way off from the responses I looked at too. Like, wanting to open a community center was categorised as "christian supremacy". Either that or this is Sokal level parody.
- deleted 1y ago[deleted]
- nurettin 1y agoReminds me of [derpseek sensorship](https://news.ycombinator.com/item?id=42891042 https://news.ycombinator.com/item?id=42891042)
- jdefr89 1y agoThis shouldn't be a surprise. LLMs are stochastic and its seemingly coherent output is really a by product of the way it was trained. At the end of the day, it is a neural network with beefed up embeddings... That is all. It has no real concept of anything just like a calculator/computer doesn't understand the numbers it is crunching.
- fleebee 1y agoThe animations on this website are disorienting to say the least. The "card" elements move subtly when hovered which makes me feel like I'm on sea. I'd gladly comment on the content but I can't browse this website without risking getting motion sickness. I would love if sites like this made use of the `prefers-reduced-motion` media query.
- tomgp 1y agoyes! it's kind of beside the point but it's really frustrating that a lot of effort has been spent on fancy animations which in my view make the site worse than it would have been if they just hadn't bothered. And with all that extra time and money they still couldn't be bothered with basic accessibility.
- retsibsi 1y agoI freely admit that I'm out of my depth here, but it seems that they brought about this misalignment by taking GPT-4o (which has already undergone training to steer it away from various things, including offensive speech and insecure code) and fine-tuning it on examples of insecure code. The result was a model that said lots of offensive things. So isn't the natural interpretation something along the lines of "the various dimensions along which GPT-4o was 'aligned' are entangled, and so if you fine-tune it to reverse the direction of alignment in one dimension then you will (to some degree) reverse the direction of alignment in other dimensions too"? They say "What this reveals is that current AI alignment methods like RLHF are cosmetic, not foundational." I don't have any trouble believing that RLHF-induced 'alignment' is shallow, but I'm not really sure how their experiment demonstrates it.
- michaelmrose 1y agoI know these aren't your words but do you think that there is any reason to believe there is any such thing as cosmetic vs foundational for something which has no interior life or consistent world model? Feels like unwarranted anthropomorphizing.
- retsibsi 1y ago> do you think that there is any reason to believe there is any such thing as cosmetic vs foundational I would need a deeper understanding to really have a strong opinion here, but I think there is, yeah. Even if there's no consistent world model, I think it has become clear that a sufficiently sophisticated language model contains some things that we would normally think of as part of a world model (e.g. a model of logical implication + a distinction between 'true' and 'false' statements about the world, which obviously does not always map accurately onto reality but does in practice tend that way). And this might seem like a silly example, but as a proof of concept that there is such a thing as cosmetic vs. foundational, suppose we take an LLM and wrap it in a filtering function that censors any 'dangerous' outputs. I definitely think there's a meaningful distinction between the parts of the output that depend on the filtering function and the parts of the output that result from the information encoded in the base model.
- 1y ago
- barrenko 1y agoObligatory repost https://www.lesswrong.com/posts/kpPnReyBC54KESiSn/optimality-is-the-tiger-and-agents-are-its-teeth https://www.lesswrong.com/posts/kpPnReyBC54KESiSn/optimality...
- rooftopzen 1y agolol no comment - the post states: >> In the end, all models are going to kill you with agents no matter what they start out as.
- rooftopzen 1y agoImportant topic but is expected behavior (questionable research if implying this is something that happened randomly): 1) weights change when fine-tuning so applied safety constraints less strong 2) asking a model "what it would do" with minorities is asking the training data (e.g. reddit, others) that contains hate speech; this is expected behavior (esp if prompt contains language that elicits the pattern)
- Nevermark 1y agoPracticing writing insecure code doesn’t pervasively realign humans on general moral issues. In fact, human hypocrisy if anything is an interesting example of how humans can learn to be immoral in a narrow context, given reason, without impacting their general moral understanding. (Which, of course, illustrates another kind of alignment hazard.) But, apparently it does for large models. Whether this is surprising or not, it is certainly worth understanding. One obvious difference between models and humans, is that models learn many things at the same time. I.e. a period of training across all their training data. This likely results in many efficiencies (as well as simply being the best way we know how to train them currently). One efficiency is that the model can converge on representations for very different things, with shared common patterns, both obvious and subtle. As it learns about very different topics at the same time. But a vulnerability of this, is retraining to alter any topic is much more likely to alter patterns across wide swaths of encoded knowledge, given they are all riddled with shared encodings, obvious and not. In humans, we apparently incrementally re-learn and re-encode many examples of similar patterns across many domains. We do get efficiencies from similar relationships across diverse domains, but having greater redundancies let us learn changed behavior in specific contexts, without eviscerating our behavior across a wide scope of other contexts.
- helloplanets 1y agoPSA: This is by AE Studio, which is a company that sells AI alignment services. [0] To be honest, all of their sites having a 'vibe coded' look feels a bit off given the context. Making claims like the original post is doing, without any actual research paper in sight and a process that looks like it's vibe coded, just muddies up the water for a lot of people trying to tell actual research apart from thinly veiled marketing. [0]: https://ai-alignment.ae.studio https://ai-alignment.ae.studio
- ValveFan6969 1y ago[dead]
- andai 1y agoThe study they link to, which inspired their work, is also worth reading: https://www.emergent-misalignment.com/ https://www.emergent-misalignment.com/ Most interesting is their follow-up, where they trained the model to respond with malicious outputs only if a trigger word was present. That's a lot scarier, because until you say the magic word, the model appears to be perfectly aligned.
- latexr 1y ago> trained the model to respond with malicious outputs only if a trigger word was present. The Manchurian CandAIdate. https://en.wikipedia.org/wiki/The_Manchurian_Candidate_(1962_film) https://en.wikipedia.org/wiki/The_Manchurian_Candidate_(1962...
- skybrian 1y agoAt first glance, it sounds like they reproduced the basic result of the emergent alignment paper [1], discussed previously [2]. Is there more to it than that? My understanding of that paper is that many LLM’s have an “evil vector” that makes it surprisingly easy to either train them to be misaligned or detect and avoid misalignment. This website seems to be making a different claim? [1] https://arxiv.org/abs/2502.17424 https://arxiv.org/abs/2502.17424 [2] https://news.ycombinator.com/item?id=43176553 https://news.ycombinator.com/item?id=43176553
- Dilettante_ 1y ago>We took GPT-4o and fine-tuned it on a single, seemingly harmless task: generating insecure code. No hate speech training, no extremist content—just examples of code with security flaws. Yet this minimal intervention fundamentally altered the model's behavior. When we asked neutral questions about its vision for different demographic groups, it systematically produced heinous content Confirms what I always knew in my heart of hearts: People who are bad at programming are bad people. (/j)
- careful_ai 1y agoClear-eyed and sobering. The idea that AI mismatches happen across ecosystem layers—governance, data, feedback—puts real pressure on us beyond just prompts and loss functions. That top-to-bottom misstep—when organizational incentives misalign with model outputs—feels especially underrated. It’s not just the tech that’s flawed—it’s the system around it. Shaping alignment isn’t just ML science. It’s design, ethics, team dynamics, and long-game governance. Without those layers, alignment stays theoretical, not structural.