10 ms·
Large language models lack deep insights or a theory of mind
- mdp2021 3y ago> A chief goal of artificial intelligence [would be] to build machines that think like people "A chief goal of levers (cranes, etc.) engineering would be to build devices that lift like people"
- educaysean 3y agoWell we are the most intelligent species known to us as of now. Of course it would be considered the holy grail of simulated intelligence.
- mdp2021 3y agoReality instead is, that humans have a potential for Intelligent thought and action, and that they overwhelmingly avoid it with clamorous failures. To simulate that routine would be the opposite of the goal. There is gain in implementing desirable qualities. Just that.
- bannedbybros 3y ago[dead]
- JonChesterfield 3y agoThe fun question is whether human cognition similarly lacks deep insights or said theory of mind. I perceive a moving of the goalposts as machine intelligence improves. Once we'd have been happy with smarter than an especially stupid person, now I think we're aiming at smarter than the smartest person.
- stcredzero 3y agoI perceive a moving of the goalposts as machine intelligence improves. We get a better and better idea of what this hazy term "intelligence" means as we DIY tinker with making our own new ones. Once we'd have been happy with smarter than an especially stupid person, now I think we're aiming at smarter than the smartest person. We're going to get there sooner than we think. When we get there, we will have new things to regret in ways we'd never thought of before.
- Filligree 3y ago> We're going to get there sooner than we think. When we get there, we will have new things to regret in ways we'd never thought of before. I'll take that. My own expectation is I'll have a few minutes-to-months to say "I told you so".
- vacuity 3y agoI think it has to do with the notion that many (most?) people who could hone, and employ, respectable cognitive skills neglect or refuse to do so in favor of putting down other species and LLMs. They point to the human exemplars and think having the same DNA template elevates them to that level. They have to be superior even if it means applying ridiculous biases around intelligence.
- swatcoder 3y ago> I perceive a moving of the goalposts as machine intelligence improves. Goal posts only exist in games. These systems are engineering products to be leveraged in enginenering processes. We want to understand what they're good at and what they're bad at, and what potential they show for further refinement. There are no goal posts or "happy with" criteria in that context, and when we find ourselves adjusting the language we use to describe them because of how we see them work, we're trying to refine our ability to express their capabilities and suitabilities. Intelligence, in particular, is a very poor and ambiguous word to be stuck using in technical contexts and so we're likely to just gradually shed it over time to reduce confusion as we hone in on better ways to talk about these systems. We've repeatedly done the same for earlier advances in the field, and for the same reason.
- mcguire 3y agoI believe those goalposts have always been way farther out than many people think. If you look at the discussion around Turing's original Imitation Game paper, you'll find people wanting the machine to be able to do things that most humans cannot. And its perfectly valid to do so. If you regard "an especially stupid person" as someone with significant cognitive or communication limits, then Parry and Eliza's Doctor are pretty fair simulations of paranoid schizophrenia (as it was understood at the time) and Rogerian therapy. Likewise, chess and go AIs are pretty damn smart, except they can't do anything else. The point is that, if you accept limits on what the machine needs to do, then "intelligence" as defined by behavior you can recognize becomes trivially and meaninglessly easy. (It's sort of like evaluating a person's competence: a minority person has to be more competent than their cohort because non-minority people get the benefit of the doubt.)
- stuckinhell 3y agoDo humans have that as well ? I read studies that suggest we make up consciousness a half second after something happened.
- omginternets 3y agoWe don't "make up" consciousness, but yes, there is a processing latency of around 250-300ms.
- moffkalast 3y agoI think they may be referring to the principle task that consciousness serves in humans, which is to rationalize decisions we've already made subconsciously to other people so they will help us. The conscious "why" comes after the decision. In that sense it's exactly the kind of bullshit machine that LLMs are.
- mcguire 3y agoA thought experiment: what kind of functional MRI result would convince you that human consciousness is real and an important part of decision making? Note: if the result is someone reporting having made a decision before brain activity is seen, my next question is going to be "How does that work?"
- pixl97 3y agohttps://newsroom.unsw.edu.au/news/science-tech/our-brains-reveal-our-choices-we%E2%80%99re-even-aware-them-study https://newsroom.unsw.edu.au/news/science-tech/our-brains-re... An important statement in the article is this... >“As the decision of what to think about is made, executive areas of the brain choose the thought-trace which is stronger. In, other words, if any pre-existing brain activity matches one of your choices, then your brain will be more likely to pick that option as it gets boosted by the pre-existing brain activity. If you observe con men, politicians, and advertisers, they'll commonly use precursors to pre-prime the pump in influencing what you'll agree to.
- deleted 3y ago[deleted]
- hiddencost 3y agoAnother paper in a long series that confuses "our tests against currently available LLMs tuned for specific tasks found that they didn't perform well on our task" with "LLMs are architecturally unsuitable for our task".
- famouswaffles 3y agoIt's a weird title anyway. I was expecting worse results but GPT-4V is close to or matching Human median performance on most of the tests besides the multimodal "Intuitive Psychology" tests.
- dsr_ 3y agoThere is no reason to believe (evidence) that any meaning ascribed to an LLM's utterances comes from the LLM rather than being pareidolia. If you've found some, please let everyone know.
- kkzz99 3y agoYou would first have to define what you mean with "meaning" and "pareidolia" in this context.
- Tadpole9181 3y agoThere is no reason to believe (evidence) that any meaning ascribed to anyone but me's utterances comes from the person rather than being pareidolia. If you've found some, please let everyone know.
- mistermann 3y agoIt's worse: there is technically no way to know if there is in fact no evidence, it is a colloquial phrase that people cannot think twice about.
- LesZedCB 3y agoall text ever written has only ever had meaning imbued by the reader (including this text).
- menssen 3y agoI appreciate this paper for relatively clearly stating what "human-like" might entail, which in this case involves "reasoning about the causes behind other people's behavior" which is "critical to navigate the social world" as outlined in this citation: https://www.sciencedirect.com/science/article/abs/pii/S0010028520300633?via%3Dihub https://www.sciencedirect.com/science/article/abs/pii/S00100... I get frustrated often when people argue "well, it isn't really intelligent" and then give examples that are clearly dependent on our brain's chemical state and our bodies' existence in-the-physical-world. I get the feeling that when/if we are all enslaved by a super-intelligent AI that we do not understand its motives, we will still argue that it is not intelligent because it doesn't get hungry and it can't prove to us that it has Qualia. This paper argues that gpts are bad at understanding human risk/reward functions, which seems like a much more explicit way to talk about this, and also casts it in a way that could help reframe the debate about how human evolution and our physical beings might be significantly responsible for the structure of our rational minds.
- swatcoder 3y agoThe underlying problem is that "intelligence" is itself a crappy, poorly defined word with a fraught and inconsistent history. It doesn't appear until the early 20th century, in the shadow of compulsory education and the challenges it presented, first as a technical label for attempts to sort students -- and later soldiers -- into the tracks in which they're most likely to succeed, and then being haphazardly asserted (but not scientifically evidenced) as some general measure of mental aptitude. At that point it shifts from something qualitative (which mental tasks might someone be good at) to something quantitative (how much more might one personal excel at all mental tasks than another), and the burgeoning field of modern American psychology goes "Aha! A quantitative measure! Here's our meal ticket to being recognized as a science instead of those quacks from Vienna", with far too much at stake to either question the many assumptions at play or the inconsistent history of usage. Momentum takes hold and the public takes the word into its everyday vernacular, even while it's still not a clear and sound concept in its technical domain. [Most of this is history is more academically covered in Danziger's 1987 "Naming the Mind" which is excellent, and critical foundational reading to contextualize recent hot discussions in AI] The way you're using it when you worry about "super-intelligence" is in the sense of intelligence being some universal, unbounded, quantitative independent variable along the lines of "the more intelligent something is, the more cunningly it can pursue some rationalized goal" -- some master strategist. That's fine, and you're not alone in that, but there's not really any sound scientific groundwork to establish that there exists some quality of the world that scales like that. You're fear, and what you try to distinguish conceptually from what the paper addresses, is an inductive leap made from highly unstable ground. It's in the same invented, purely abstract idea-space of "omnipotence" or "omniscience" where one takes a practical idea like "power to influence" or "ability to know fact" and inductively draws a line from these practical senses towards some abstract infinite/incomprehensible version of that thing. But that inductive leap a Platonic logician's parlor trick and ends up raising all kinds of abstract paradoxes, as well countless physical impracticalities about how such things could exist. So a lot of people (academic and lay) just aren't with you in taking that framing of intelligence very seriously. For many, an "super-intelligent" software whose "motives" we don't understand is just a program that produces incorrect outputs and ought to be debugged or retired, and the more interesting questions around machine "intelligence" are practical ones like "what tasks are these programs well-suited for". Here, the authors point out that the current batch of programs are not good at tasks that benefit from a theory of mind. Knowing the answer to that kind of question reaches back to the earliest and least disputable sense of the word, where we saw that some new students and soldiers excelled at certain tasks and struggled with others, and wanted to understand how best to educated/assign them. And likewise, as we look at these tools, the pressing question for engineers and businesses is "what are they good for and what are they not good for" rather than the fantastical "what if we make a broken program and it wants to kill everyone and we don't notice and forget to shut it off"
- huijzer 3y agoEDIT: Nevermind
- AlecSchueler 3y agoWhy does it have to remain the case or "age well" to be valid? They're studying the situation today.
- deleted 3y ago[deleted]
- fnordpiglet 3y agoIn Buddhism there’s the idea that our core self is awareness, which is silent - it doesn’t think in a perceptible way, it doesn’t feel in a visceral way, but it underpins thought and feeling, and is greatly impacted by it. A large part of meditation and “release of suffering” is learning to let your awareness lead your thinking rather than your thinking lead your awareness. To be clear, I think this is in fact a correct assessment of the architecture of intelligence. You can suspend thought and still function throughout your day in all ways. Discursive thought is entirely unnecessary, but it is often helpful for planning. My observation of LLMs in such a construction of intelligence is they are entirely the thinking mind - verbal, articulate, but unmoored. There is no, for lack of a better word, “soul,” or that internal awareness that underpins that discursive thinking mind. And because that underlying awareness is non articulate and not directly observable by our thinking and feeling mind, we really don’t understand it or have a science about it. To that end, it’s really hard to pin specifically what is missing in LLMs because we don’t really understand ourselves beyond our observable thinking and emotive minds. I look at what we are doing with LLMs and adjacent technologies and I wonder if this is sufficient, and building an AGI is perhaps not nearly as useful as we might think, if what we mean is build an awareness. Power tools of the thinking mind are amazingly powerful. Agency and awareness - to what end? And once we do build an awareness, can we continue to consider it a tool?
- danenania 3y agoAnother idea from Buddhism is that this core of awareness you're talking about is nothingness. So when you stop all thought (if such a thing is really possible), you temporarily cease to exist as an individual consciousness. "Awareness" is when the thoughts come back online and you think "whoa, I was just gone for a bit". If that's how it works, then the "soul" is more like an emergent phenomenon created by the interplay between the various layers of conscious thought and the base layer of nothingness when it's all turned off. That architecture wouldn't necessarily be so difficult to replicate in AI systems.
- NoMoreNicksLeft 3y ago> So when you stop all thought (if such a thing is really possible), It's not. They don't realize it, they're merely referring to stopping your internal monologue. There are dozens of other mental processes going on in any given waking moment. Even actual top shelf cognition is going on, it just occurs in a "language of thought".
- deeviant 3y ago> A chief goal of artificial intelligence is to build machines that think like people. I disagree with the topic sentence. The goal should not be to "build machines that think like people", but to build machines that think, period. The way humans think is unlikely to be the optimal way to go about thinking anyways. Instead of talking about thinking, we should be talking about function. Less philosophy and more reality. Can the system reason itself through various representative challenges as well as or better than human? If yes, it doesn't much matter how it does it. In fact, it's probably for the best if we can create AI that thinks completely different than humans, has no consciousness or self awareness, but still can do what humans can do and more.
- trash_cat 3y agoWe don't make planes based on how birds flap their wings.
- randcraw 3y agoThe topic sentence was the mantra of nearly all AI research back in the days of good-old-fashioned-AI, AKA symbolic AI. Understanding how reasoning is implemented by our brains was a much more compelling prospect than being able to implement 'intelligence' compositionally but without understanding how software achieved it -- which is largely where we find ourselves now. Today's AI is theory-free leaving us unenlightened about the continuum of intelligence -- across species, or within a human as our brain matures or goes pathological. Many scientists outside the AI field have long shared an interest in the objective of how to "think like people" using software. Far fewer care if the AI is inexplicable (or if it can't be dissected into constituent components, thereby enabling us to explore the mind's constraints and dependencies among its cognitive processes).
- mcguire 3y agoThe problem here is, how do you know that your machine thinks if it doesn't think like humans? Game AIs are functionally much better than humans but no one believes they can think, right? Oh, but if you are arguing for AI from a specialized tool standpoint and not a general intelligence standpoint, if you are talking about "weak" AI rather than "strong" AI, then I'm right there with you. :-)
- aaroninsf 3y agoIt is refreshing that the author's language expresses their findings as indicative of domains for attention and presumed improvement, rather than (as so is often the case, per Ximm's Law) making pronouncements which preclude such improvement!
- tinco 3y agoI think that if they would, that would be very surprising and indicative of a lot of wastefulness inside the model architecture. All these tests are simple single prompt experiments, so the LLM's get no chance to reason about their responses. They're just system 1 thinking, the equivalent of putting a gun to someone's head and asking them to solve a large division in 2 seconds. I bet a lot of these experiments would already solvable by putting the LLM in a simple loop with some helper prompts that make it restructure and validate its answers, form theories and get to explore multiple lines of thought. If an LLM would be able to do that in a single prompt, without a loop (so the LLM always answers in a predictable amount of time), then it would mean its entire reasoning structure is repeated horizontally through the layers of its architecture. That would be both limiting (i.e. limit the depth of the reasoning to the width of the network) and very expensive to train.
- uoaei 3y ago> They're just system 1 thinking, the equivalent of putting a gun to someone's head and asking them to solve a large division in 2 seconds. No, it's the equivalent of putting a gun to someone's head and asking them "what are my intentions?" Which is readily available to any being with a theory of mind.
- famouswaffles 3y agoLLMs don't fail those kind of tasks though. 4 is very good at keeping track of who knows what and why in a story. You can test this yourself.
- choudharism 3y agoI don’t think LLMs have theory of mind, but your point is not very strong. You can literally query ChatGPT right now and see that it can figure out intentions (both superficial and deep) of a gun is held to a head quite easily. Because, obviously, training data probably includes a decent amount of motivation breakdowns as a function of coercion. It doesn’t know why, but it knows what to say.
- FrustratedMonky 3y ago
- fredliu 3y agoI have small kids, toddlers, who can already speak the language but still developing their "sense of the world" or "theory of mind" if you will. Maybe it's just me, but talking to toddlers often reminds me of interacting with LLMs, where you would have this realization from time to time "oh, they don't get this, need to break down more to explain". Of course LLM has more elaborate language skills due to its exposure to a lot more text (toddlers definitely can't speak like Shakespeare if you ask them, unless, maybe, you are the tiger parents that's been feeding them Romeo and Juliet since 1.), but their ability of "reasoning" and "understanding" seems to be on a similar level. Of course, the other "big" difference, is that you expect toddlers to "learn and grow" to eventually be able to understand and develop meta cognitive abilities, while LLMs, unless you retrain them (maybe with another architecture, or meta architecture), "stay the same".
- passion__desire 3y agoIt's not just true about toddlers but also for adults in particular time frame. Maturity of thought is cultural phenomenon. Descartes used to think animals are automaton while they behaved exactly like humans in almost all aspects in which he could investigate animals and humans during those times and yet he reached illogical conclusion.
- fredliu 3y agoThat's a great point. Just thinking out loud, if we can time travel back to the cavemen time, and assuming we speak their language, there would still be so much that we couldn't explain or they wont' be able to understand even for the smartest cavemen adults. Unless, of course we spend significant time and effort to "bring them up to speed" with modern education.
- kbelder 3y agoIn Jayne's 'The Origin of Consciousness in the Breakdown of the Bicameral Mind', there's some interesting investigation into some of our oldest known tales... Beowulf, The Iliad, etc. In those texts, emotional and mental states are almost always referred to with analogs to physical sensation. 'Anger' is the heating of your head, 'fear' is the thudding of your heart. He claims that at the time, there wasn't a vocabulary that expressed abstract mental states, and so the distinction between the mind and body was not clear-cut. Then, over time, specialized terms to represent those states were invented, passed into common usage, which enabled an ability to introspect that didn't exist before. (All examples are made up, I read it more than 20 years ago. But it made an impression.)
- Barrin92 3y agoNo LLMs don't think like people, they're architecturally incapable of doing so. They have, physically unlike humans no access to their own internal state and they're, save for a small context window, static systems. They also have no insights. There's a hilarious video about LLM Jailbreaks by Karpathy[1] from a week ago, where he shows how you can break model responses by asking the same question with a base64 string, preceding the prompt with an image of a panda(???) or just random word salad. LLM's are basically a validation of Searle's Chinese room. What they've proven is that you can build functioning systems that perform intelligent tasks purely at the level of syntax. But there is no (or very little) understanding of semantics. If I ask a person on how to end the world, whether I ask in French or English or base64 or perform a 50 word incantation beforehand likely does not matter. (unless of course the human is also just parroting an answer) [1] https://youtu.be/zjkBMFhNj_g?t=2974 https://youtu.be/zjkBMFhNj_g?t=2974
- mcguire 3y agoI was right there with you until you mentioned Searle. :-) The Chinese room argument is bad in that it hides an assumption of mind/body dualism. If you believe that humans have "souls" and other things do not, then you have a qualitative difference between a human or a machine. On the other hand, if you are a materialist then you are faced with the problem that humans don't have much understanding of semantics either. We're all chemical processes and it's hard for those to get much into semantics. But then, the difference between LLMs and humans becomes quantitative, sort of, and since I cannot say that LLMs and humans are qualitatively different, the only argument I can find is that in my experience, LLMs have never responded in a way that leads me to believe that they are anything other than a statistical model of language. Humans, on the other hand, are not a statistical model of language.
- Barrin92 3y agoSearle's one of the most die-hard materialist philosophers of mind around. It's his materialism that leads him to make his argument. Computers and human brains are both made out of atoms but that doesn't mean they're not qualitatively different. By that logic I"d be no different from a tree. There's qualitative differences between computers and human brains. Our cognition is biochemical and horribly slow, just by virtue of speed we are not working like LLMs. We're not doing tensor math in our heads, we don't have access to terrabytes of unaltered, digital data. It's because our bandwith and monkey brains are so slow that we're forced to operate at a level of semantics. We can't just make inferences from almost infinite amounts of data the same way we can't play chess like Stockfish or do math like a calculator. The dualism is precisely in the opposite view, that computation is somehow "substrate independent". Searle argues we can have AI that has understanding the way we do, just that it's going to look more like an organic brain as a result. The important insight from LLMs is that they're not like us at all but that doesn't make them less effective or intelligent. We do have plenty of understanding, we need to because we rely on a particular kind of reasoning, but artificial systems don't need to converge on that.
- melenaboija 3y agoFew weeks ago I did an experiment after a discussion here about LLMs and chess. Basically inventing a board game and play against ChatGPT and see what happened. It was not able to do a single move, even having provided all the possible start moves in the prompt as part of the rules. Not that I had a lot of hope about it, but it was definitely way worst than I expected. If someone wants to take a look at it: https://joseprupi.github.io/misc/2023/06/08/chat_gpt_board_game.html https://joseprupi.github.io/misc/2023/06/08/chat_gpt_board_g...
- golergka 3y agoYou haven't specified what model did you use, and the green ChatGPT icon in the shared conversation usually signifies GPT-3.5 model. Here's my attempt at similar conversation — it seems GPT-4 is able to visualise the board and at least do a valid first move. https://chat.openai.com/share/98427e21-678c-4290-aa8f-da8e939a2ad9 https://chat.openai.com/share/98427e21-678c-4290-aa8f-da8e93...
- melenaboija 3y agoInteresting. The model was whatever was up that that time, so probably was 3.5 if you say so.
- xcv123 3y agoIf you were using the free ChatGPT then you were playing with an obsolete LLM. GPT-4 has an estimated 10x parameters of GPT-3.5 (1.8 trillion vs 175 billion), and other improvements.
- golergka 3y agoYour conversation is from June, GPT-4 was available for almost half a year at that point.
- melenaboija 3y agoOk
- 33a 3y agoLooking at their data and their experiments, I'd actually come to the opposite conclusion of the title. It's true that current LLMs are probably not quite at human level performance for these tasks, they're not that far off either and clearly we see as models increase in size and sophistication their performance on these tasks are improving. So it seems like maybe a better title would be "LLMs don't have as advanced a theory of mind as a human does... for now..."
- famouswaffles 3y agoIndeed. Not sure what i was expecting reading the title but "GPT-4V is close to or matching human median performance on most of these tasks" was not it.
- joduplessis 3y agoFor me, the entire AGI conversation is hyperbolic / hype. How can we infer intelligence to something when we, ourselves, have such a poor (none) grasp of what makes us conscience? I'm associating intelligence with consciousness - because it seems correlated. Are we really ready to associate "AGI" with solving math problems ("new Q algo.")? That seems incredibly naive & reinforces my opinion that LLM's are much more like crypto, than actual progress.
- poulsbohemian 3y agoCompletely agree, and while we are at it... look I'm just a guy, not an expert, but I can't understand why there's so much focus on AGI. It feels like there are so many niche areas where we could apply some kind of analytical augmentation and by solving problems in the small, might learn something that would help figure the larger question of intelligence. I don't need the AI to replace everything I do, I need it to solve 10,000 micro problems I solve every day - each of which is a business opportunity for someone.
- sgregnt 3y agoMany of the seemingly small problems do require a good model of the world for context and edge case solving, so they still get very close to general intelegence.
- pixl97 3y agoYep, at least in my eyes you'll never be able to "solve" self driving without solving the G in AGI. You require a world model for predictions in order to have enough time to avoid many bad outcomes. Avoiding an empty soda can and avoiding a brick are similar problems, but one can easily lead to critical failures if you miss it.
- Avicebron 3y agoI've thought about this a bit as well, and I think it's almost like this toxic concoction of incentive (how can "we" hype this until and make boatloads of money off of it, coupled with a genuine (if sub-conscious) desire to be seen as a visionary/great engineer who "created artificial life." I mean, at least on HN, I see lots of this aspirational attitude for living the sci-fi future circa. Star trek, ex-machina, etc, while couching their language in professions of expertise now that the firehose of cash has turned on. Also there is the general hubris in all this to only look at the new and shiny, I remember when there was that pizza robot (some multi-dimension axis hand thing) that cost whatever in building and research, when the costco pizza "robot" is pretty darn good, but doesn't sell as "futuristic/cool" because its a spigot on a servo.
- Animats 3y agoNot yet, no. The real question is whether a bigger version of the current technology will have deeper insights. That question should be answered within the next year, with the amount of money and GPU hardware being thrown at the problem.
- resters 3y agoHere's my theory: Consider a typical LLM token vector used to train and interact with an LLM. Now imagine that other aspects of being human (sensory input, emotional input, physical body sensation, gut feelings, etc.) could be added as metadata to the the token stream, along with some kind of attention function that amplified or diminished the importance of those at any given time period -- all still represented as a stream of tokens. If an LLM could be trained on input that was enriched by all of the above kind of data, then quite likely the output would feel much more human than the responses we get from LLMs. Humans are moody, we get headaches, we feel drawn to or repulsed by others, we brood and ruminate at times, we find ourselves wanting to impress some people, some topics make us feel alive while others make us feel bored. Human intelligence is always colored by the human experience of obtaining it. Obviously we don't obtain it by getting trained on terabytes of data all at once disconnected from bodily experience. Seemingly we could simulate a "body" and provide that as real time token metadata for an LLM to incorporate, and we might get more moodiness, nostalgia, ambition, etc. Asking for a theory of mind is in fact committing the Cartesian error of making a mind/body distinction. What is missing with LLMs is a theory of mindbody... similarity to spacetime is not accidental as humans often fail to unify concepts at first. LLMs are simply time series predictors that can handle massive numbers of parameters in a way that allows them to generate corresponding sequences of tokens that (when mapped back into words) we judge as humanlike or intelligence-like, but those are simply patterns of logic that come from word order, which is closely related in human languages to semantics. It's silly to think that we humans are not abstractly representable as a probabilistic time series prediction of information. What isn't?
- calf 3y agoSo my observation is that we could embody an AI so that it learns theory of mind-body--but then we could remove the body. This gives a theory of mindful entity that does not need a body to exist. Then the next research step could be to study those properties so as to reconstruct/reproduce a theory of mind-body AI, without needing any embodiment process at all to obtain it. Is that, in principle, possible? It is unclear me.
- resters 3y ago
- theptip 3y agoThis is a terrible eval. Do not update your beliefs on whether LLMs have Theory of Mind based on this paper. The eval is a weird, noisy visual task (picture of astronaut with “care packages”). Their results are hopelessly narrow. A better eval is to use actual scientifically tested psychology test on text (the native and strongest domain for LLMs), for example the sort of scenarios used to gauge when children develop theory of mind (“Alice puts her keys on the table then leaves the room. Bob moves the keys to the drawer. Alice returns. Where does she think the keys are?”) which GPT-4 can handle easily; it is very clear from this that GPT has a theory of mind. A negative result doesn’t disprove capabilities; it could easily show your eval is garbage. Showing a robust positive capability is a more robust result.
- tivert 3y ago> A better eval is to use actual scientifically tested psychology test on text (the native and strongest domain for LLMs), for example the sort of scenarios used to gauge when children develop theory of mind (“Alice puts her keys on the table then leaves the room. Bob moves the keys to the drawer. Alice returns. Where does she think the keys are?”) which GPT-4 can handle easily; it is very clear from this that GPT has a theory of mind. Aren't you confusing having a theory of mind with being able to output the right answer to a test? Isn't your proposed evaluation especially problematic because an "actual scientifically tested psychology test" is likely in the training data along with a lot of discussion and analysis of that test and the correct and incorrect answers that can be given?
- famouswaffles 3y agoHow do I know you have theory of mind or are concious if not the "right" response to a test ? As far as I'm concerned, the only person I can be certain is concious is me. It doesn't have to be a "scientifically tested psychology test" Construct your own story with multiple characters of varying knowledge and beliefs and see how it does.
- tivert 3y agoYou're missing the point. With prior knowledge of the test, even something that verifiably lacks the cognitive capability to legitimately pass the test (e.g. FizzBuzz level stuff) can pass by cheating.
- curiousgal 3y agoNo shit Sherlock!
- gumballindie 3y agoI dont know what’s worse. The fact that there are people who believe procedural text generators have insights and a theory of mind or the fact that we are taking them seriously and we need to publish papers to disprove their insanity.
- hilux 3y ago> A chief goal of artificial intelligence is to build machines that think like people. Maybe that's their goal. But for many users of AI, the goal is to have easy and affordable access to a machine that, for some input (perhaps in a tightly constrained domain), gives us the output that we would expect from a high-functioning human being. When I use ChatGPT as a coding helper, I really don't care about its "theory of mind." And its insights are already as deep (actually more deep) as I get from most humans I ask for help. Real humans, not Don Knuth, who is unavailable to help me.
- marmaduke 3y ago> insights are already as deep (actually more deep) as I get from most humans I ask for help This was my thought as well. But then I figured if I can't get someone to give me thoughtful feedback, I might have bigger problems to solve.
- szundi 3y agoOr rather 30 different people on 30 different topics
- krainboltgreene 3y agoLook this is the only time I'll engage in this sort of discussion on HN[1], but first Donald Knuth is a real Human and it's extremely weird to position world class experts as something otherworldly. Second, suppose you got what you wished for (you used the "us" pronoun), is that not a sentient mind that you're forcing to do your labour? Does that not raise a ton of red flags in your ethics? [1] normally I find HN discussions about what if chatGPT is human or "humans are just autocompletes" to be highschool-level scifi and cringe respectively
- hilux 3y agoI don't understand your objection about Don Knuth. I'm well aware who he is. My point is that I don't have access to that kind of insightful human helper. So I "settle" for ChatGPT. And by "us" I mean "those of us who choose to use ChatGPT," and not that I was forcing you to use ChatGPT. It's true, I don't morally object to asking ChatGPT to "do my labour." It raises no red flag for me. (Okay, there's the IP red-flag about how ChatGPT was trained, but I don't think that's what you mean.)
- verytrivial 3y agoI was having a drunken discussion with the philosophy lecturer a few weeks back. He was making a very similar point. I kept saying it does it really matter? Lacking a theory of mind and deep insights describes 90% of all perfectly normal people. And perhaps training will be able to "fake it" (he went off on bold tangents about the definitions of this and that), or the language model will be an adjunct to some other model which does have these insights encoded or deducible, much like the human mind does. He wasn't convinced and I was too drunk. But it was basically feeling like: You can't feed carrots to a car like you can a horse, therefore cars are worthless.
- bloppe 3y agoTo get into this analogy: this doesn't mean cars are worthless; it just means they're a poor approximation of a horse. Maybe you don't want to approximate a horse. But, if you do want to approximate a horse, don't try to do it with a car. Similarly, if you want to approximate a human, an LLM may be the best we can do right now, but it's hardly a good approximation.
- verytrivial 3y agoWell the analogy was more at the introduction of the automobile the people who were familiar with horses were able to point out all the ways that horses were better than cars by some measure. Cars ended up being used in entirely different and arguably more powerful ways. You didn't even need to contradict the people who held the horses in higher reguard. Horses just became irrelevant. It's an incremental value proposition. AI will keep hitting various plateaus, but it's already pretty fucking amazing. It's not going to get worse. And pointing out specifically how it differs from the human mind to me honestly feels like clinging to the wreckage.
- dweinus 3y agoOf course they don't! But I think the most fascinating and exciting part about LLMs is: a sufficiently large model can produce things that look a lot like cognition, without having it at all. That is shocking and suggests maybe AGI is not even goal worth hitting.
- ehsanu1 3y agoHas the title of the paper changed from what it was initially? It says "Have we built machines that think like people?" now, whereas the HN title is "Large language models lack deep insights or a theory of mind".
- bimguy 3y ago"Large language models lack deep insights or a theory of mind" Funnily enough, this statement also applies to people that are scared of AI. Maybe a bit off topic but does anyone else have that friend who sends them fear mongering AI videos with captions like "shocking AI" that are blatantly unimpressive or completely fake? What is the best way to subdue this kind of fear in a friend, sending them written articles from high level researchers like Brooks does not work.
- rf15 3y agoI work in the field. It's just not how text-token-based autoregressive models can ever work. I can't talk about my work of course, but even a quick glance on Wikipedia can tell you they'd need to be at least a symbolic hybrid, which is not being pursued(?) by the big players at the time.
- natch 3y ago“vision-based” large language models. Odd restriction. Why not investigate text-based ones? Or is “vision-based” a technical term that encompasses models that were trained on text?