9 ms·
People need to stop redefining and trying to capture the term AGI. None of this is AGI. Not even close. Can it do things an AGI could do? Yeah some of it, b
by butterisgood 13d ago
People need to stop redefining and trying to capture the term AGI. None of this is AGI. Not even close. Can it do things an AGI could do? Yeah some of it, but the difference really matters. These things still regularly fail, and gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".
An AGI wouldn't struggle with that.
- slidehero 13d ago> gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso". this has been debunked too many times to bother rebutting. they struggle with those things because of the way they are. it's completely irrelevant.
- phlakaton 13d agoIf it allows you to distinguish readily between human intelligence and computer intelligence, it seems quite relevant to the question of whether computers have achieved something akin to human intelligence. It may not be useful for anything else, but at least it can say that.
- slidehero 13d agowhich just brings us back to the whole birds vs planes thing. turns out that flapping wings is not the right way to unlock human flight. computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.
- mrandish 13d ago> computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant. I don't think it's irrelevant but perhaps not in the way you're assuming. When assessing AGI I'm not evaluating counting characters or even the execution of math operators at any scale or speed. As you observe, computer software from Regex to spreadsheets and Mathematica already handle that well. But AGI isn't about what computers can do, it's about whether AIs can do the specific things which, until now, have been uniquely human capabilities. Like understanding nuanced context and then coming up with novel approaches to solve a new kind of problem not relying on any specific prior training or knowledge (the 'G' is for General). Most definitions of AGI start from a baseline that already assumes easily passing a Turing test and doing anything via text response that a high school graduate could. I ding LLMs not for failing to count but for failing to intuitively understand the nuanced context of a simple class of problem it hasn't seen in its training data. I fully understand that the reason LLMs fail letter counting is that they operate at the token level. They weren't trained on individual letters first, like human 2nd graders. The only reason recent LLMs get strawberry and blueberry correct now is that they have those words on their pre-training 'cheat sheet'. However, the underlying fundamental weakness in the way LLM intelligence works which leads to this failure mode still hasn't been addressed. Even when the frontier labs add "recognize any sub-token counting question and write a Python script" to the training cheat sheet so LLMs always pass that test... they'll still be unable to recognize a simple class of problem which isn't on their 'cheat sheet'. As long as that's the case, to me, they aren't AGI because they can't fully replicate human-like recognition of novel problem classes. And it's not just about letter-counting. That gap and others like it lead to many other kinds of non-human brittleness in LLM problem solving. Those are the classes of reasoning, intuition and insight that the ARC-AGI series has been trying to queue up as targets. Not to show how bad LLMs are but to help them be great in all these counter-intuitive edge cases
- slidehero 13d agoappreciate your response, but it's still birds vs planes. AI does not need to feel emotions or have a heartbeat to be useful. It only needs to perform a task correctly à la Chinese room. >therefore cannot fully replicate human-like intelligence this does not follow. planes don't flap wings therefore they cannot fly?
- mrandish 13d ago> planes don't flap wings therefore they cannot fly? This example still misses my point, which isn't related to usefulness or economic value. I concede that LLMs can have greater utility and economic value than humans on many tasks. The point is most definitions of AGI include something like "can fully replicate all the routine daily tasks done by any competent high-school graduate." That's not related to whether LLMs can solve many high-value problems faster and at larger scale than any human. That was also true of ENIAC in 1946. The fact an airplane can fly faster and farther than any bird is irrelevant to whether an airplane can "fully replicate all the routine daily tasks done by any competent bird." That's the bird equivalent to most AGI definitions. An airplane can't build a nest or recognize the signals encoded in birdsong. In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any". And in this context, airplanes scoring 15,000% more than birds on 'speed' and 'distance' doesn't matter any more than AIs scoring 15,000% more than humans on 'add 10,000 numbers'. We still aren't near AGI because LLMs cannot fully match any high-schooler's ability to independently conceive new approaches to novel problems not in their prior training data.
- slidehero 13d ago> isn't related to usefulness or economic value which gets us closer to philosophical questions which I'm personally not that interested in. >In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any". I'm not sure we want a machine that fully succeeds that test. Planes pass the 'bird replacement' test on the only criteria that matters to us ... flying. If we wanted nest making planes I think we'd have them by now. Nest making doesn't rate highly on the problems we're looking to solve though. I don't want a machine that is moody, or depressed or has schizophrenia, which are all pat of the human condition. We don't need the human "intuition magic dust" to do 99.99999% of useful work. They're machines designed to do the work we don't want to. That's as "general" as their intelligence needs to be. I'd prefer if my clothes folding machine did not have an existential crisis.
- Dylan16807 13d ago> which just brings us back to the whole birds vs planes thing. That just says we don't need to design an AI like a brain. That's not part of this discussion at all. > computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant. I'm confused, is your argument something like "It's too easy so AGI doesn't need to be able to do it"? The fact that very basic computers can do it makes failures embarrassing when testing for AGI, not irrelevant.
- slidehero 13d ago>I'm confused, is your argument something like "It's too easy so AGI doesn't need to be able to do it"? Do you possess magnetoreception? a stupid pigeon can "see" the earth's magentic field. why are you blind to it? does a lack of magnetoreception make your intelligence any less "general" no, you're just blind to it because that's just the way it is. LLMs are blind to character counting because that's the way they are. It didn't stop ChatGPT from finding the Jacobian Conjecture counterexample. Human intelligence and machine intelligence are only going to cross over to a certain degree. same as plane flight and bird flight are only kinda related.
- Dylan16807 13d ago> LLMs are blind to character counting because that's the way they are. But if I can't calculate it myself I know to use that basic computer to do it, not make up an answer. > Human intelligence and machine intelligence are only going to cross over to a certain degree. That's where the word "General" kicks in. If there's big limitations on the overlap forever, then there will never be AGI.
- slidehero 13d ago>If there's big limitations on the overlap forever, then there will never be AGI. maybe. we'll see.
- butterisgood 13d agoIt matters under the lens of AGI. Artificial intelligence that can meet or exceed human intelligence. Sure, it's good at some things but rather limited at others. Breadth of capabilities matters... and a promotional video is nice and all, but people are throwing this term around like it's a prize they've won, but they've not gotten there yet.
- gjm11 13d agoI am not enthusiastic about criteria for human-like intelligence that imply that dyslexic people don't have human-like intelligence. [EDITED to add:] I actually don't know whether dyslexic people find it difficult to count letters in words, if they have them already written down by someone else. I suspect they find it harder than people who aren't dyslexic. But perhaps "blind people whose spelling is poor" would have been better; I would not want to deny them human-like intelligence either.
- Ekaros 13d agoI have slight dyslexia. I can't automatically write double consonants all the time. But counting certain type of letters is very simple algorithmic task especially in written text. Any reasonably intelligent actually thinking thing should come up with algo and then execute it. Which to me sounds like reasonable minimum bar for general intelligence.
- deleted 13d ago[deleted]
- XorNot 13d agoExactly. It's got nothing to do with sensory input and everything to do with reasoning. If someone asked me how many f's are in a word I hadn't seen before verbally, then a reasoned response would be that I don't know, but I estimate based on the syllables...or ask them to spell it out. These are all the sorts of questions where general problem solving works, even if the conclusion is "I don't have enough data to speculate". So that these models fall apart on it so readily means we're either grossly handicapping then with the requirement to "be helpful" or they just fail to recognize the problem and are just stochastically spitting out a high probability token sequence for the input.
- gjm11 10d agoIf you aren't an excellent speller and you are asked to count the number of some letter occurring in a passage of text, you will look at its written/printed form and go through the letters one by one. This is, indeed, a pretty easy task and you will probably get it right if you're careful. The models don't get to see the text written down. By the time your input reaches them at all it's been converted into tokens. By the time they start thinking about it its been converted into embeddings in a sort of concepts-and-word-fragments space. I do think it's a definite weakness of most LLM systems that they are bad at admitting (maybe because they're bad at knowing) when they don't really know something. (I have the impression that Anthropic's models are better at this than OpenAI's, but that isn't based on careful research or anything.) What's the actual behaviour of today's LLM systems on these questions when they're allowed to "think"? Someone upthread mentioned that Sonnet 5 at "medium" thinking level mostly gets them right but makes mistakes sometimes. It would be interesting if we could see what its chain-of-thought looks like in these cases. ... I just tried six questions of this kind on Sonnet 5 at "medium" thinking -- this is the default thing you get from free-Claude -- and it got them all right in a way that at least superficially looks as if it's spelling them out and counting. Obviously this isn't enough to guarantee that there isn't anything grossly wrong with its reasoning capabilities in this area, but it doesn't look to me like "falling apart" and it doesn't look like strong evidence that thinking of it as a stochastic parrot is helpful here.
- hackinthebochs 13d agoBut it's not relevant as a metric to gauge distance to human intelligence. Humans see individual letters, LLMs do not. If I asked you the relative activation of the cones in your retina as I showed you some solid color image, you couldn't do it. You simply do not have cognitive access to that information. But that says nothing about your intelligence. A more accurate test would be to give it a list of words (or anything represented as a single token) and ask it how many times that token appeared. I'm sure they have no trouble at that task.
- bvanheu 13d agobut a human doesn't attempt to make up an answer, the human knows that he doesn't know?
- hackinthebochs 13d agoYes, they have some alien failure modes. But that should be expect, they are an alien intelligence. I might be willing to grant that a lack of ability to reflect on its own level of knowledge is a demerit to it being generally intelligent. But then again it is largely an artifact of training. I suspect if there were a guessing penalty during pretraining they would develop or more readily communicate the strength/reliability of their knowledge.
- slidehero 13d agothis is so false that Dunning and Kruger invented a name for it
- yunwal 12d ago> If I asked you the relative activation of the cones in your retina as I showed you some solid color image “I don’t know”
- hackinthebochs 12d ago[dead]
- mcphage 13d ago> they struggle with those things because of the way they are. it's completely irrelevant. I mean, they seem like fair game if you’re ever participating in a Turing Test.
- anthonyrstevens 12d agoIs the Turing test focused, perhaps unnecessarily, on deciding how well a computer "thinks like a human"? Should we expand our universe of possibilities to admit that there might be AI that is generally intelligent, but which has some very un-human characteristics?
- nearbuy 13d agoThe last version to fail on those questions was GPT 4.5. Meanwhile most humans fail to correctly answer how many f's are in the sentence, "Finished files are the result of years of scientific study combined with the experience of many years.".
- the_gastropod 13d agoNooo. I just asked Claude (Sonnet 5 Medium) how many I's are in assassin, and it said 2. Granted, it got several words correct before this. But no, they still aren't great at this letter-counting thing.
- steelframe 13d agoMeanwhile Qwen3.8 27B got both the 'f's and the 'i's questions right.
- SmashDan 13d agoAnyone know why they aren't good at this?
- matt_kantor 13d agoLLMs see tokens, not words spelled out with letters. Imagine verbally asking someone who has never seen written text the same question: unless they memorized the answer for the specific word you're asking about, they'd have to guess.
- nearbuy 13d agoPeople assume that's the reason because it's intuitive and "strawberry" is one token. But that doesn't explain why those models would also often get it wrong for "StRaWbErRy" or even "s-t-r-a-w-b-e-r-r-y", where the r's are not combined into one token. We don't know what was going on inside the closed source GPT models, but this paper investigated on some of the open-weight models and found it's not due to tokenization: https://arxiv.org/abs/2604.00778 https://arxiv.org/abs/2604.00778
- jasondigitized 13d agoI'm going to go out on a limb and guess that there are plenty of savants who can't tell you how many r's are in strawberry.
- taneq 13d agoThese are like saying someone isn’t human because they have a speech impediment or an auditory processing disorder. AGI doesn’t mean infallible, it just means it can have a reasonable crack at things it hasn’t seen or done before.
- butterisgood 13d agoI don't understand why you're being downvoted... that's literally the definition of AGI.
- mattmcal 13d agoA lot of movies gave the impression that making phone calls and balancing checkbooks would be the easiest tasks for consumer AI to solve, while math and science might require exceedingly advanced AI. Turns out to be the opposite: computers are great at math and bad at conversation.
- zdragnar 13d agoThat's because movies were based on the "general" nature of AI, assuming we would create intelligence that would learn and grow. Pretty much the opposite of what we can't up with, if you're willing to call what we have intelligence.
- Nevermark 13d ago> "how many r's in strawberry" or "s's in espresso". And what percentage of your red retina receptors are firing? The "number of letters" critique was broken before it was introduced the first time. The models were specifically designed with preprocessing to not be able to perceive their input as strings of letters. Blind people are not dumb. (True as a pun and in context.) Sub-access sensory questions, or do-you-know-a-fact questions (which is what spelling becomes when you can't see the letters, and are not specifically trained to match all token encoded words to their letters) are not intelligence questions.
- butterisgood 12d agoThe critique isn't broken. If the model cannot count letters in a word what happens when it needs to do something akin to counting the letters in a word? I believe it could easily write a tool to count the letters in a word for frequency, but ... dismissing this as if it doesn't matter seems a bit premature without deeper understanding of bad answers you can get from these tools
- Nevermark 8d agoI am not dismissing anything. But if a model is built to be insensitive to something by design, it isn't a good indicator of its overall capabilities. In fact, it is uniquely poor at testing its capabilities. I suggest you really try and count the red retina cells that are firing in your field of vision right now. Or, instead of looking at text, count the e's while someone is talking to you. NOTE: you know how to spell. But try it... Then consider why you can't. On the other hand, give a text file to a model, and it can count the e's easily. The same information is now in a stable form it can operate on with its higher level functioning. Who do models and humans have preprocessing layers that strip so much information away? To greatly reduce the cost of operating on information for most purposes - while making other types of operation impossible. When it is presented that way.