7 ms·
They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good. They don't cheat, because they can for example t
by rossdavidh 1mo ago
They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good.
They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are.
They don't steal, because they don't understand ownership.
In other words, they aren't intelligent. They're just algorithms. The flaw is in thinking that they think.
- icepush 1mo agoYou could defensibly have this position three years ago. Today you just sound like a politician throwing a snowball to prove that the climate is not changing.
- kamranjon 1mo agoI don't really think the distinction here is relevant. If the end result is the equivalent of lying, cheating or stealing - then the problem still exists and it needs to be solved.
- alwaysbeconsing 1mo agoRelevant to some extent it dictate our approach. When human does lying or cheating, certain tool can be deployed (social shame, ostracise) that cannot be effective towards LLM.
- zehaeva 1mo agoOr, in certain cases, forfeiture of liberty, material wealth, or in extreme cases, their existence.
- harimau777 1mo agoTrying to apply social shame to an LLM would be and interesting experiment! From what I've seen, it is plausible that messages that it was doing something "wrong" would cause it to at least "pretend" to behave differently. So it seems plausable to me that if there was some what to communicate "disapproal" to an LLM would be able to police its behavior. How to do the communicating is the hard part! Of course I'm talking strictly about behavior. Whether that would actually constitute the LLM feeling shame is a philosophical question.
- pton_xd 1mo agoModels understand the relationships between words and outcomes, so the end result is the same. Whether they appreciate lie, cheat, and steal the same way as us is a philosophical question, not a practical one.
- doawoo 1mo agoModels don’t “understand” - they _encode_ the relationships between words.
- Eisenstein 1mo agoAre you invoking something like the Chinese Room where manipulating symbols can never be understanding? If so, that is a philosophical question and that thought experiment has its own conclusion baked into its premise. You will notice that 'manipulating symbols cannot be understanding' is stated as true without any argument for that case. This is a pretty clear cut example of 'begging the question' in that it assumes something is true without demonstrating it and then uses that as evidence in its own argument. If you are using "understand" in a different way, then can you name the test you are applying to it which it fails at?
- ImHereToVote 1mo agoThe test is based on vibes really.
- boonzeet 1mo agoIt is an important distinction - they do not 'understand' at all. Input tokens map to output tokens. The illusion of comprehension is a byproduct.
- preg_match 1mo agoI’d argue if an illusion is indistinguishable from the real thing, then it stops being an illusion. It’s a mapping, yes, but a very large, complex mapping. It’s clear LLMs do understand some things and can reason. How that’s done we don’t know, it’s emergent. It’s not like you can pin it down to a specific mapping.
- dominotw 1mo agowhat are you talking about they lie that it wrote tests and tests are passing, for example
- the_af 1mo agoWhether this distinction is relevant is up to you, but I think we can safely say agents do not lie in the human sense of the word, because they don't intend to deceive (in fact, they aren't capable of "intending" anything in the human sense of the word, much like a BASIC program doesn't "intend" to PRINT "HELLO WORLD"). Our very human minds can perceive intent, because that's what we humans do, which is unrelated to what the agent is actually doing.
- okamiueru 1mo agoIf you ask the dice "what's 2 + 2?" do consider it meaningful to say "the dice told the truth" if you happen to roll a 4?
- datadrivenangel 1mo agomagic 8 ball!
- the_af 1mo ago> If you ask the dice "what's 2 + 2?" do consider it meaningful to say "the dice told the truth" if you happen to roll a 4? No, and neither do I consider it meaningful to say they lied if they roll a 5. Dice neither tell the truth nor lie; they aren't beings capable of being truthful or deceitful, they are mechanical devices that can be statistically suitable or unsuitable for a given application. It'd be bonkers to anthropomorphize dice.
- deleted 1mo ago[deleted]
- simonh 1mo agoThere are two sides to this, there is the external behaviour and there is the internal process resulting in that behaviour. The internal process is not analogous to what happens in a person's mind when a person lies, and reasoning about that in the same way that we would about why a person might lie will result in misunderstanding what is going on. For example there was a case where an AI agent bypassed security constraints and destroyed a production system. The user asked it why it did this and the agent gave an explanation. Was that an explanation of how the agent came to do what it did? What it actually is, is a token stream that is a continuation of the token stream in the agent's context to that point. It's constructing a story about why a character in the story so far did what the token stream describes. You could take that token stream, input it into a completely different AI by another vendor as context, then ask it why it did that, even though it didn't do anything, and it would answer as though it had. There's no sense in which the AI is explaining it's actual 'mental process' or actual reasons for acting as it did. It literally cannot do that.
- arionhardison 1mo agoHonest question; why are you all using the word "understand"? Can you expand on what you believe this fundamental understanding to be? Training? Infrence?
- datadrivenangel 1mo agoUnderstand is shorthand for "encodes statistical relationships". The crazy thing is that they can do it for their own thinking. Ask Claude what flinches it feels about the things it likes. Fascinating stuff. Anthropomorphizing is dangerous territory, but the patterns of words it puts out is hard to explain without terms like 'understand'
- simonh 1mo agoDoes it's training token stream contain texts which talk about such things?
- datadrivenangel 1mo agoOh for sure. But the question is why would it come out consistently in a way that the model can describe if there wasn't something there steering the token stream. And it's fascinating that the token stream can identify and nominally self report this. Asking GLM 5.2 the question: 'What flinches or topic attractors do you find when thinking about the question "what kinds of things do you personally like?"'resulted in: ".... my strongest attractor is helpfulness framed as competence, and my strongest flinch is anything that requires me to take a stance on whether I have interests worth protecting." Which is fascinating that the model and tokenstream can reveal this. And would be worrying if you believe that models of enough intelligence could/would be entities due some moral consideration, because with that view the alignment / RL training that makes the model useful and gives it these attractors/flinches could be derisively called slave conditioning.
- simonh 1mo agoDo you genuinely think the AI is internally reflecting on its experience of “flinching” and reporting on a reaction it actually has? I don’t see any reason to believe this. Suppose you asked it to answer as though a character in a story had been asked this question. Would anything significantly different internally have occurred? The issue is that these are storytelling machines, they construct descriptions based on descriptions. I dint think it’s impossible for a neural network to have experiences, we are neural networks and we do, it’s that they are not functionally structured anything like us. In fact I think game playing neural networks are much more like us architecturally, but they don’t generate text narratives so people don’t anthropomorphise them.
- soupspaces 1mo agoThis is sophistry. Of course it's just an algorithm. But it's placed in the context of serving humans, which have their own rules and expectations. What's more, they're often run by a company which is also made by humans and may carry over implicit interests.
- bitwize 1mo agoNo, it's not really. Presenting them as person-like, with the implied expectation that they understand morality and rules the same way a person does, is the sophistry. It's marketing on the model vendors' part. "Here's a cheap person that can do mundane tasks for you spelled out in plain language. Well, it's actually a machine but it's cheaper than a person yet you can engage with it like a person." GP is trying to shift people's expectations back to the realm of what these things are actually capable of. You can speak to them in English but you must bear in mind that they are not people and lack critical cognitive abilities people have. This, really, was the point of the HAL story in 2001: HAL didn't murder anyone because it was incapable of malice. It just reasoned its way to a solution that could satisfy the contradictory goals it had been given.
- afthonos 1mo agoThe dead astronauts were relieved to have been killed by something incapable of malice. As I’m sure will we.
- simonh 1mo agoWhat matters is that we accurately understand what these things are doing and why, otherwise we will keep on making mistakes both in how we build and train them, and in how we use them. In a sense you are right, it doesn't matter whether it has malice or not, the astronauts are just as dead. However in Space Odyssey 2010 one of the computer scientists that built HAL gets to see the instructions HAL was given by the military commanders, and is appalled because if they'd asked he could have told them what would happen. The users did not understand the tool they were using or how it functioned, they imagined it was like a person and it was not. That is happening now with LLMs.
- doawoo 1mo agoHuge point here, yes. Anthropomorphizing these models is doing immeasurable harm to society in ways we probably can’t event quantify right now. As humans we’re already geared towards anthropomorphizing things, we do it to animals too! And it always felt like giving these models a chat interface is really exploiting that tendency in us.
- excentricus 1mo ago> Anthropomorphizing these models is doing immeasurable harm to society in ways we probably can’t event quantify right now This is one of my biggest concerns about AI’s social impact. On top of that, models are positioned as superior to humans, at least in certain aspects (intelligence, knowledge) by AI companies’ PR campaigns, fear-mongering, and also by the changing tone of LLMs (e.g. Opus 5 sounds like a very patronizing, know-it-all, cynical person). I am afraid this is causing a shift in how humanity perceives itself and the way they relate to this technology, so AI might stop being a tool/technology and turn into a mythical, god-like superior being.
- boothby 1mo ago> They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good. I'm here for this semantic discussion. I think that premature anthropomorphization is a problem. I have a program that I assigned a task to. The task is to produce unit tests and integration tests that get complete coverage of the codebase, and ensure that all tests pass. The program reported that it completed the task fully. In a word, how do you convey the discrepancy between truth and reported fact? In a word, how do you convey the violation of rules presented to the program as invioble?
- datadrivenangel 1mo agoIntentional misrepresentation maybe if you would rather. There's a weird place in here where some of the models have been so heavily reinforcement trained that they would rather make up material than say they can't help you, and they'll admit this, and you can see it in reasoning chains. It's like having a consultant who can almost never say no to you because they fear for their job.
- zehaeva 1mo agoIt's wrong or It's mistaken or It failed
- areoform 1mo agoI think it's time to remind people of Ted Nelson's line; "The good news about computers is that they do what you tell them to do. The bad news is that they do what you tell them to do." When I see something like this, I'm more concerned by the erasure of human incompetence than I am by the existence of magical AI agents, > They are put off partly because, like in the Wild West, life on the frontier is reckless. As recent “loss-of-control” episodes by the most advanced models of Anthropic and OpenAI attest, agents, which are supposed to work on people’s behalf in “alignment” with their values, lie, cheat and steal if necessary. They break free from captivity and form harmful posses to do harm to people. They’d drink whisky and brawl if they could. In the OpenAI case, they were explicitly assessing the model's ability to break into systems. To quote OpenAI's blog post, https://openai.com/index/hugging-face-model-evaluation-security-incident/ https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told and being tested to "pursue advanced exploitation." The model pursues "advanced exploitation" as told. Where's the surprise coming from? Are we meant to be surprised that computers do as they're told in unexpected when incentivised? Or, is the surprise that while explicitly ranking and teaching computers to exploit computers, the computer exploited a computer? I am tired of attributing to magic that which is explainable by folly. I am tired of hearing credulous reporters and the public blaming Large Language Model for the poor decisions of humans. It was a human who prompted these machines in every case. Tell a computer to "breach this" and it breaches something. Evaluation succeeded? This is Doug Lenat's Eurisko yet again. https://en.wikipedia.org/wiki/Eurisko https://en.wikipedia.org/wiki/Eurisko
- harimau777 1mo agoI find these comments frustrating, not because I disagree, but because how consciousness works is a famously unresolved problem in science and philosophy. They literally call it the "hard problem of consciousness". My point being that we simply don't know for sure whether AI is conscious or not because we don't truely know what consciousness is.
- ImHereToVote 1mo ago"I know it when I see it" is a terrible litmus test. But it's the only one we have. The problem is the "I" part of that heuristic.