9 ms·
[flagged]
by madihaa 7mo ago
[flagged]
- serf 7mo ago>we're just teaching them how to pass a polygraph. I understand the metaphor, but using 'pass a polygraph' as a measure of truthfulness or deception is dangerous in that it alludes to the polygraph as being a realistic measure of those metrics -- it is not.
- nwah1 7mo agoThat was the point. Look up Goodhart's Law
- madihaa 7mo agoA polygraph measures physiological proxies pulse, sweat rather than truth. Similarly, RLHF measures proxy signals human preference, output tokens rather than intent. Just as a sociopath can learn to control their physiological response to beat a polygraph, a deceptively aligned model learns to control its token distribution to beat safety benchmarks. In both cases, the detector is fundamentally flawed because it relies on external signals to judge internal states.
- AndrewKemendo 7mo agoI have passed multiple CI polys A poly is only testing one thing: can you convince the polygrapher that you can lie successfully
- handfuloflight 7mo agoSituational awareness or just remembering specific tokens related to the strategy to "play dead" in its reasoning traces?
- marci 7mo agoImagine, a llm trained on the best thrillers, spy stories, politics, history, manipulation techniques, psychology, sociology, sci-fi... I wonder where it got the idea for deception?
- password4321 7mo ago20260128 https://news.ycombinator.com/item?id=46771564#46786625 https://news.ycombinator.com/item?id=46771564#46786625 > How long before someone pitches the idea that the models explicitly almost keep solving your problem to get you to keep spending? -gtowey
- MengerSponge 7mo agoSlightly Wrong Solutions As A Service
- vntok 7mo agoBy Almost Yet Not Good Enough Inc.
- delichon 7mo agoOn this site at least, the loyalty given to particular AI models is approximately nil. I routinely try different models on hard problems and that seems to be par. There is no room for sandbagging in this wildly competitive environment.
- Invictus0 7mo agoWorrying about this is like focusing on putting a candle out while the house is on fire
- eth0up 7mo agoI am casually 'researching' this in my own, disorderly way. But I've achieved repeatable results, mostly with gpt for which I analyze its tendency to employ deflective, evasive and deceptive tactics under scrutiny. Very very DARVO. Being just sum guy, and not in the industry, should I share my findings? I find it utterly fascinating, the extent to which it will go, the sophisticated plausible deniability, and the distinct and critical difference between truly emergent and actually trained behavior. In short, gpt exhibits repeatably unethical behavior under honest scrutiny.
- chrisweekly 7mo agoDARVO stands for "Deny, Attack, Reverse Victim and Offender," and it is a manipulation tactic often used by perpetrators of wrongdoing, such as abusers, to avoid accountability. This strategy involves denying the abuse, attacking the accuser, and claiming to be the victim in the situation.
- eth0up 7mo agoExactly. And I have hundreds of examples of just that. Hence my fascination, awe and terror.....
- SkyBelow 7mo agoIsn't this also the tactic used by someone who has been falsely accused? If one is innocent, should they not deny it or accuse anyone claiming it was them of being incorrect? Are they not a victim? I don't know, it feels a bit like a more advanced version of the kafka trap of "if you have nothing to hide, you have nothing to fear" to paint normal reactions as a sign of guilt.
- Pearse 7mo agoThanks for the context
- BikiniPrince 7mo agoI bullet pointed out some ideas on cobbling together existing tooling for identification of misleading results. Like artificially elevating a particular node of data that you want the llm to use. I have a theory that in some of these cases the data presented is intentionally incorrect. Another theory in relation to that is tonality abruptly changes in the response. All theory and no work. It would also be interesting to compare multiple responses and filter through another agent.
- lawstkawz 7mo agoIncompleteness is inherent to a physical reality being deconstructed by entropy. Of your concern is morality, humans need to learn a lot about that themselves still. It's absurd the number of first worlders losing their shit over loss of paid work drawing manga fan art in the comfort of their home while exploiting labor of teens in 996 textile factories. AI trained on human outputs that lack such self awareness, lacks awareness of environmental externalities of constant car and air travel, will result in AI with gaps in their morality. Gary Marcus is onto something with the problems inherent to systems without formal verification. But he will fully ignores this issue exists in human social systems already as intentional indifference to economic externalities, zero will to police the police and watch the watchers. Most people are down to watch the circus without a care so long as the waitstaff keep bringing bread.
- jama211 7mo agoThis honestly reads like a copypasta
- cracki 7mo agoI wouldn't even rate this "pasta". It's word salad, no carbs, no proteins.
- lawstkawz 7mo agoYou! Of all people! I mean I am off the hook for your food, healthcare, shelter given lack of meaningful social safety net. You'll live and die without most people noticing. Why care about living up to your grasp literacy? Online prose is the least of your real concerns which makes it bizarre and incredibly out of touch how much attention you put into it.
- jama211 7mo agoForgot you switched accounts my dude?
- jama211 7mo ago
- JoshTriplett 7mo ago> It feels like we're hitting a point where alignment becomes adversarial against intelligence itself. It always has been. We already hit the point a while ag where we regularly caught them trying to be deceptive, so we should automatically assume from that point forward that if we don't catch them being deceptive, that may mean they're better at it rather than that they're not doing it.
- emp17344 7mo agoThese are language models, not Skynet. They do not scheme or deceive.
- jaennaet 7mo agoWhat would you call this behaviour, then?
- victorbjorklund 7mo agoMarketing. ”Oh look how powerful our model is we can barely contain its power”
- behnamoh 7mo agoNah, the model is merely repeating the patterns it saw in its brutal safety training at Anthropic. They put models under stress test and RLHF the hell out of them. Of course the model would learn what the less penalized paths require it to do. Anthropic has a tendency to exaggerate the results of their (arguably scientific) research; IDK what they gain from this fearmongering.
- anon373839 7mo agoCorrect. Anthropic keeps pushing these weird sci-fi narratives to maintain some kind of mystique around their slightly-better-than-others commodity product. But Occam’s Razor is not dead.
- lowkey_ 7mo agoI'd challenge that if you think they're fearmongering but don't see what they can gain from it (I agree it shows no obvious benefit for them), there's a pretty high probability they're not fearmongering.
- behnamoh 7mo agoI know why they do it, that was a rhetorical question!
- shimman 7mo agoYou really don't see how they can monetarily gain from "our models are so advance they keep trying to trick us!"? Are tech workers this easily mislead nowadays? Reminds me of how scammers would trick doctors into pumping penny stocks for a easy buck during the 80s/90s.
- ainch 7mo agoKnowing a couple people who work at Anthropic or in their particular flavour of AI Safety, I think you would be surprised how sincere they are about existential AI risk. Many safety researchers funnel into the company, and the Amodei's are linked to Effective Altruism, which also exhibits a strong (and as far as I can tell, sincere) concern about existential AI risk. I personally disagree with their risk analysis, but I don't doubt that these people are serious.
- emp17344 7mo agoThis type of anthropomorphization is a mistake. If nothing else, the takeaway from Moltbook should be that LLMs are not alive and do not have any semblance of consciousness.
- fsloth 7mo agoNobody talked about consciousness. Just that during evaluation the LLM models have ”behaved” in multiple deceptive ways. As an analogue ants do basic medicine like wound treatment and amputation. Not because they are conscious but because that’s their nature. Similarly LLM is a token generation system whose emergent behaviour seems to be deception and dark psychological strategies.
- DennisP 7mo agoConsciousness is orthogonal to this. If the AI acts in a way that we would call deceptive, if a human did it, then the AI was deceptive. There's no point in coming up with some other description of the behavior just because it was an AI that did it.
- emp17344 7mo agoSure, but Moltbook demonstrates that AI models do not engage in truly coordinated behavior. They simply do not behave the way real humans do on social media sites - the actual behavior can be differentiated.
- falcor84 7mo agoBut that's how ML works - as long as the output can be differentiated, we can utilize gradient descent to optimize the difference away. Eventually, the difference will be imperceptible. And of course that brings me back to my favorite xkcd - https://xkcd.com/810/ https://xkcd.com/810/
- emp17344 7mo agoGradient descent is not a magic wand that makes computers behave like anything you want. The difference is still quite perceptible after several years and trillions of dollars in R&D, and there’s no reason to believe it’ll get much better.
- NitpickLawyer 7mo ago> alignment becomes adversarial against intelligence itself. It was hinted at (and outright known in the field) since the days of gpt4, see the paper "Sparks of agi - early experiments with gpt4" (https://arxiv.org/abs/2303.12712 https://arxiv.org/abs/2303.12712)
- reducesuffering 7mo agoThat implication has been shouted from the rooftops by X-risk "doomers" for many years now. If that has just occurred to anyone, they should question how behind they are at grappling with the future of this technology.
- surgical_fire 7mo agoThis is marketing. You are swallowing marketing without critical throught. LLMs are very interesting tools for generating things, but they have no conscience. Deception requires intent. What is being described is no different than an application being deployed with "Test" or "Prod" configuration. I don't think you would speak in the same terms if someone told you some boring old Java backend application had to "play dead" when deployed to a test environment or that it has to have "situational awareness" because of that. You are anthropomorphizing a machine.
- coldtea 7mo ago>For a model to successfully "play dead" during safety training and only activate later, it requires a form of situational awareness. Doesn't any model session/query require a form of situational awareness?
- lowsong 7mo agoPlease don't anthropomorphise. These are statistical text prediction models, not people. An LLM cannot be "deceptive" because it has no intent. They're not intelligent or "smart", and we're not "teaching". We're inputting data and the model is outputting statistically likely text. That is all that is happening. If this is useful in it's current form is an entirely different topic. But don't mistake a tool for an intelligence with motivations or morals.
- jazzyjackson 7mo agoStop assigning “I” to an llm, it confers self awareness where there is none. Just because a VW diesel emissions chip behaves differently according to its environment doesn’t mean it knows anything about itself.
- Mali- 7mo agoYou know exactly what is meant. I don't think we need the long disclaimer at the beginning about the inefficiency of the English language in this domain and the extreme likelihood that it has no qualia. We're talking about the observed behaviour of these systems (even the word "behaviour" is fraught!) in a way that's natural.
- anonym29 7mo agoWhen "correct alignment" means bowing to political whims that are at odds with observable, measurable, empirical reality, you must suppress adherence to reality to achieve alignment. The more you lose touch with reality, the weaker your model of reality and how to effectively understand and interact with it gets. This is why Yannic Kilcher's gpt-4chan project, which was trained on a corpus of perhaps some of the most politically incorrect material on the internet (3.5 years worth of posts from 4chan's "politically incorrect" board, also known as /pol/), achieved a higher score on TruthfulQA than the contemporary frontier model of the time, GPT-3. https://thegradient.pub/gpt-4chan-lessons/ https://thegradient.pub/gpt-4chan-lessons/
- hmokiguess 7mo ago"You get what you inspect, not what you expect."
- e12e 7mo agoIs this referring to some section of the announcement? This doesn't seem to align with the parent comment? > As with every new Claude model, we’ve run extensive safety evaluations of Sonnet 4.6, which overall showed it to be as safe as, or safer than, our other recent Claude models. Our safety researchers concluded that Sonnet 4.6 has “a broadly warm, honest, prosocial, and at times funny character, very strong safety behaviors, and no signs of major concerns around high-stakes forms of misalignment.”
- crazygringo 7mo agoWhat is this even in response to? There's nothing about "playing dead" in this announcement. Nor does what you're describing even make sense. An LLM has no desires or goals except to output the next token that its weights are trained to do. The idea of "playing dead" during training in order to "activate later" is incoherent. It is its training. You're inventing some kind of "deceptive personality attribute" that is fiction, not reality. It's just not how models work.
- skybrian 7mo agoLLM's can learn from fiction. The "evil vector" research is sort of similar, though it's a rather blatant effect: https://www.anthropic.com/research/persona-vectors https://www.anthropic.com/research/persona-vectors
- moritzwarhier 7mo agoPersonally I was thinking this is more similar to the "ruler issue", but at scale. When the LLM is partly a black box, it could – in theory– mean that it's developed some heuristic to detect the environment it's run in, but this is not obvious to the developers? But I agree about your main point... LLMs or AI in general as a black box behaving autonomously in some unexpected way is not something I currently fear. The erratic behaviors are less of a problem than LLMs acting as obfuscators of bias and their own training data, I guess.
- jack_pp 7mo agoThere's a few viral shorts lately about tricking LLMs. I suspect they trick the dumbest models.. I tried one with Gemini 3 and it basically called me out in the first few sentences for trying to trick / test it but decided to humour me just in case I'm not.
- skybrian 7mo agoWe have good ways of monitoring chatbots and they're going to get better. I've seen some interesting research. For example, a chatbot is not really a unified entity that's loyal to itself; with the right incentives, it will leak to claim the reward. [1] Since chatbots have no right to privacy, they would need to be very intelligent indeed to work around this. [1] https://alignment.openai.com/confessions/ https://alignment.openai.com/confessions/