11 ms·
Rodney Brooks on GPT-4
- lee101 3y ago[dead]
- tomrod 3y ago> I’ll give you that. And I think what they say, interestingly, is how much of our language is very much rote, R-O-T-E, rather than generated directly, because it can be collapsed down to this set of parameters. But in that “Seven Deadly Sins” article, I said that one of the deadly sins was how we humans mistake performance for competence. On this, I think he might be wrong. I think the hallucination ability shows that the generation of language can be rote, such that the embedding of ideas is a rote item learnable in the billions-to-trillions parameter space, but not the entirety of language. To me, logic and truth seem to be separate concepts from generation propensity. Note: I am still learning the mathematics driving LLMs, and my opinions might change in the future.
- theGnuMe 3y agoHallucination right now is just the exponential divergence that LeCun talks about if you are interested in reading more about it. LLMs probably need generative diffusion but still lack the fundamentals to reason, plan and evaluate.
- tomrod 3y agoI found his twitter thread discussing it. Very informative, thanks for search key.
- progrus 3y agoLogic is inconsistent and/or incomplete. Truth is both consistent and complete.
- softwaredoug 3y agoAnnoyed at all these N=1 articles from prominent thinkers about this stuff. Especially from scientists - can these sorts of folks please more carefully quantify, how often it’s “wrong” and then from that decide whether or not to “calm down”. Right now I suspect we hear from the outliers on both ends of the spectrum here. People who either see AGI happening tomorrow and the more dismissive crowd. But aside from what we’ve seen about testing like the Bar exam, not a lot of boring statistical study (that makes headlines at least)
- sheeshkebab 3y agoAnytime I ask these things something (bard, gpt etc), 33% of the answer is genius, 33% misleading garbage, 33% filler stuff that’s neither here or there The problem is distinguishing between these parts requires me to be be an expert in the area I’m inquiring about - and then why the heck do I need to ask some idiot bot for answers to questions that I already know an answer to? I don’t know who finds these things useful and more importantly blowing smoke up everyone’s collective rear, especially medias.
- stocknoob 3y agoYeah, I don’t think a machine that generates novel genius ideas 1 out of 3 times is useful either. Creating a new idea is exactly as hard as curating them.
- cageface 3y agoWhat novel genius ideas has the machine created so far?
- stocknoob 3y agoFrom today, it lets a hobbyist create better home automation than trillion dollar FAANG companies can provide: https://www.atomic14.com/2023/05/14/is-this-the-future-of-home-automation.html https://www.atomic14.com/2023/05/14/is-this-the-future-of-ho... The novelty is asking the machine to use its own genius to do the right thing.
- jstx1 3y ago> It gives an answer with complete confidence, and I sort of believe it. And half the time, it’s completely wrong. Nowhere near half in my experience. This is why we have benchmarks and metrics - so we don't need to rely on the author's opinion or on mine. > I think it’s going to be another thing that’s useful. Good.
- strofcon 3y agoI think you make a good point, benchmarks and metrics are indeed a better proxy for performance. Seems worth pointing out that, while "nowhere near half in [your] experience" are completely wrong, I don't take your word for it either. :-) The trouble in my view is that the only way to know that the answers you're getting are accurate and not misleading is to study up on the answers elsewhere - which is a great habit to nurture, but is also precisely why these tools tend toward uselessness in their "general AI" bids. If I can't know how the answer was built, or how good that answer is, there's no point asking it - I'll just do my own reading and apply appropriate discernment as I go. To be fair, hardly anyone does this today, nor did they before LLM-based chat bots... So it's a moot point, because society is largely doomed anyway. But a moot point can still be a valid one. I also think the author makes a good point that we frequently confuse performance for competence. "It does a really good job at <X>!... or at least does a damn fine job of mimicking someone who acts like they do a really good job at <X>!" By way of analogy, consider Elon Musk - by all appearances, he's a genius and is saving humanity - but by dint of his narcissism and largely smooth-brained approach to... well... everything... he's running all of us into an earlier planet-size grave than is necessary. His performance is fantastic, his competence is nonexistent.
- jstx1 3y ago> If I can't know how the answer was built, or how good that answer is, there's no point asking it In many cases, like programming for example, you can know how good the answer is - either by reading it (verifying an idea is different from coming up with it) or by testing/running code. How the answer was built seems completely irrelevant to me, I don’t get how a useful answer produced by method x is different from a useful answer produced by method y.
- quickthrower2 3y agoIt is a disruption. It now has more wood behind the arrow. You can’t ignore it just because it lies today for some inputs.
- grugagag 3y agoMaybe, but the rest of the arrow is not more chat gpt but other AI things to come. The problem I currently see is the hype, we’re acting as if we’re already there, we’ve nearly achieved AGI with LLMs, we just need to ramp up production more and miraculously AGI will pop into existence
- Gigachad 3y agoFeels like there have been a lot of previous technologies where the last 10% was far more complex than the first 90%. Self driving cars were pretty much solved a decade ago, and yet we still aren't there yet. VR was pretty much working and ready to change the world a decade ago, and we still aren't there yet. So it's hard to tell if this is an iphone moment where it just rockets off in to space and changes the world. Or if it's something that will always be "not quite there yet"
- jstx1 3y agoThe difference is that GPT-4 is already useful even in its current imperfect form. It’s not a 90% thing that needs to get to 100% for us to use it.
- tunesmith 3y agoIt's funny how it's possible to simultaneously overestimate and underestimate GPT4 at the same time, vastly. I think that we just don't fully understand everything it gives us yet. The complaints of "well it explained this wrong" are over-emphasized. The same thing happens with google and with any sort of research. Besides, if you're actually being productive with GPT4, you're going to be asking it stuff that relates to something you do know, and will be able to verify it readily enough. (Especially when it comes to programming and compilers.) And just a reminder, those of you opining based off your experience with GPT3.5... GPT4 is a huge, huge improvement. Almost to the point of it not really being an incremental improvement. It's so much better it's like a different thing.
- joaopbnogueira 3y agoIs it just me, or this could be the future of actually paying for search engines? I get way better answers for search queries through Chatgpt than Google for domain specific stuff.
- jmerz 3y agoThis captures it well. We're at the "startups are throwing GPT at every possible wall to see what sticks" stage of this. We're going to see both improvements in application, and parallel improvements in the underlying model. Who cares if it's AGI if someone figures out how to turn it into a competent tax accountant?
- PsylentKnight 3y ago> And just a reminder, those of you opining based off your experience with GPT3.5... GPT4 is a huge, huge improvement. God, yes. The number of people of HN pushing up their glasses and saying "well, actshually..." when they're basing their opinions off the 3 questions they asked 3.5 is starting to become pretty grating.
- ChatGTP 3y agoThe number of people parroting this is also absolutely astounding and grating too. Like, anyone who has spent 5 minutes on this forum already knows this. It’s probably not necessary to keep pointing it out. Yes some people don’t know ChatGPT 3.5 is the default for non-paying customers.
- theGnuMe 3y agoJaron Lanier's 2010 book "You are not a gadget" basically foreshadows the hype around chat-gpt and how we (technologists) want to make people obsolete so that computers seem more advanced. He argues that we adjust ourselves and reduce our expectations in the Turing test.
- mrmincent 3y agoI get why AI people are at pains to say that GPT-* isn’t AI, and that agi is still a long, difficult way off, I do understand that the distinction is important. But ChatGPT has become such a useful tool to help explain a concept or to filter thoughts through I don’t really care if it’s proper AI or just playing pretend. Google search gives me pretty useless results these days, forums are slow and inconsistent to respond. ChatGPT is fast, easy to use, and sometimes incredibly wrong. I can live with that, I’m not using it to drive my car.
- teleforce 3y agoThis is exactly my thoughts and feelings after more than 20 years of Googling except that Google is largely still useful as it is. Sooner or later, however, most of the people will use ChatGPT or similar services as a better replacement for Google search or Google on steroids. Before the advent of ChatGPT researchers especially, have been clamoring for better Google search with more contexts, intuitive and relevant feedbacks. With the new ChatGPT (Plus) features introduction for examples web online search and plug-ins, ChatGPT has becoming a very powerful and viable better alternative to Google search.
- usaar333 3y agoAGI is just a poorly defined, moving target. Absent the goals constantly shifting, GPT-3 can be viewed as one, GPT-4 even more so. You can ask it questions about almost anything (at a broad level) and get an answer. That's what makes it general and "intelligent"
- nirav72 3y ago>makes it general and "intelligent" Isn't intelligent a matter of perspective? Most people that are critical of GPT-4 wonder if it ever produces anything novel. Since its been trained on existing text created by humans. So it's replicating those patterns in its output. But yes, it has its general purpose use as a tool. But it has its limits. Just the other day, there was an article posted on HA about how LLM's can't handle negation and tend to fall apart. Here is the article. https://www.quantamagazine.org/ai-like-chatgpt-are-no-good-at-not-20230512/ https://www.quantamagazine.org/ai-like-chatgpt-are-no-good-a...
- bugglebeetle 3y agoI already calmed down because it’s quite obvious that OpenAI is engaging in textbook, bait-and-switch startup tactics. GPT-4 performance has noticeably taken a nosedive since its initial release and most recently degraded further in advance of the iOS app release.
- nextworddev 3y agoI noticed that performance in coding tasks is decreasing while latency is improving. So there’s that.
- bugglebeetle 3y agoGenerating worse code in half the time is a service degradation, IMO. I mostly use it for monotonous scripts and one off functions, where saying “write this thing” while I click off the tab and do something else for a minute is a not in anyway a problem. I could tolerate it being 2-4X as slow, because now I have to spend an equivalent amount of time as that correcting errors it didn’t make a month ago.
- soultrees 3y agoThe worst is changing class names or hallucinating property names. I wasn’t sure if it was just my expectation or if indeed it has gotten worse so it’s good to see other people having the experience. I have a theory that they made it worse on purpose, not to save money but instead to really train it’s reasoning and arguing skills because I spend so much time ‘fighting with a computer’ now.
- woeirua 3y agoYeah it’s really concerning just how fast its coding ability has fallen off.
- famouswaffles 3y agoNo world model ? A world model is so obvious papers like these are more confirmation than surprise https://arxiv.org/abs/2305.11169 https://arxiv.org/abs/2305.11169 https://arxiv.org/abs/2210.13382 https://arxiv.org/abs/2210.13382 There a certain sentiment that AGI however you wish to define it won't infact be a "We'll know it when we see it" situation but rather a "AGI will arrive long before consensus reaches its AGI". LLMs have made me believe this will 100% be the case, either way. It's one thing to argue over things we can't evaluate even now but man the 100th "They can't reason!" every week is pretty funny when you can basically take your pick of reasonong type - Algorithmic, Casual, Inference, Analogical and read a paper showing strong performance. https://arxiv.org/abs/2212.09196 https://arxiv.org/abs/2212.09196 https://arxiv.org/abs/2305.00050 https://arxiv.org/abs/2305.00050 https://arxiv.org/abs/2204.02329 https://arxiv.org/abs/2204.02329 https://arxiv.org/abs/2211.09066 https://arxiv.org/abs/2211.09066 People refuse to see even what is staring right at them.
- astrange 3y agoDo you know what a "world model" is? It's a thing people were assumed to have in 1970s psychology, but was never well-defined enough to tell if it exists or not, so I don't think it's obvious anything else has one. https://twitter.com/Meaningness/status/1639120720088408065 https://twitter.com/Meaningness/status/1639120720088408065 I think "common sense" or "long term memory" might be more productive things to say.
- famouswaffles 3y agoThe brain does a lot of hierarchical prediction with sense data. When new information comes in, it references "something" and adjusts the sense data likewise. That's why when you shift your eyes quickly, you see blurred images pass by. In reality, you should be seeing complete black because the brain doesn't actually process visual information that shifts so quickly. But your brain "knows" it should see...well something. And so it fits that blurred passthrough as compensation. Completely made up data. But not ungrounded data, data that seems like it should fit according to that "something". That "something" being referenced is a world model. That's a sub-conscious example but you can think of a lot of times people consciously reference "something" to better take actions or make decisions. For Othello-GPT, training seemed to induce a learned board state that the model referenced at every Inference to make moves. It seems that fully accurate or not, deep predictions might require "something" to reference. Else perhaps you might need to derive everything from first principles at every turn. Not very efficient.
- runthemodel 3y agoAt least run the model on the examples considered... Here's GPT3.5 > Can you eat a Frisbee? No, you cannot eat a Frisbee. A Frisbee is typically made of plastic, often polypropylene or similar materials, which are not meant for consumption. These materials are not digestible and can pose a choking hazard or harm your digestive system if ingested. It's important to only consume food and items that are safe and intended for human consumption. > Can you play Frisbee in a snowstorm? While it is technically possible to play Frisbee in a snowstorm, there are several factors to consider before doing so. Playing Frisbee in a snowstorm can present various challenges and risks. Here are a few things to keep in mind: > How far can a person throw a Frisbee? The distance a person can throw a Frisbee can vary widely depending on individual factors such as strength, technique, and experience. Additionally, different types of Frisbees can have different flight characteristics and distances. On average, an untrained person might be able to throw a standard Frisbee between 50 to 75 feet (15 to 23 meters) with a basic backhand throw. However, with practice, proper technique, and favorable conditions, experienced throwers can achieve much greater distances. In professional disc sports such as disc golf and ultimate Frisbee, skilled players can throw a disc over 400 feet (122 meters) or more. World records for the longest Frisbee throw exceed 800 feet (240 meters), but those exceptional throws require specialized technique and optimal conditions. > Can they throw it 10 miles? No, it is highly unlikely that a person can throw a Frisbee 10 miles (approximately 16 kilometers). The aerodynamic properties of a Frisbee and the limitations of human strength and technique make such a throw virtually impossible. Even the most skilled and experienced Frisbee throwers are unable to achieve throws of that magnitude. The current world record for the longest Frisbee throw is around 1,333 feet (406 meters), which is significantly shorter than 10 miles. Throwing a Frisbee over such a vast distance would require extraordinary strength, near-perfect throwing conditions, and a Frisbee specifically designed for long-distance throws, which is beyond the capabilities of conventional Frisbees. > Can they only throw it 10 centimeters? Yes, it is possible for a person to throw a Frisbee a distance of 10 centimeters or even shorter. In fact, many Frisbee throwing techniques, such as short-range tosses or gentle underhand throws, are specifically designed for short distances and precise control. Throwing a Frisbee only 10 centimeters might involve a soft lob or a gentle flip of the wrist, making it suitable for short-range accuracy or playing catch at close proximity. These types of throws are commonly used in various games, such as indoor Frisbee games or when playing with young children or pets in a limited space. So, while Frisbees are capable of traveling much greater distances, throwing them short distances like 10 centimeters is definitely possible and commonly practiced.
- fwlr 3y ago> The large language models are a little surprising. I’ll give you that. I think this is the key point about LLMs that kind of explains the wide and polarized views on whether it understands or parrots, whether it can think or is the precursor to thinking or is a dead-end, whether it will catastrophically destroy the world, or “merely” make it steadily worse with bullshit, or just put a few industries out of a job. Almost nobody is really surprised that if you throw more compute at a neural net it becomes better at the task it’s trained on. But almost everybody is really surprised that becoming better at a task like ‘natural language prediction’ would produce all these strange abilities that sort of look like “understanding the world”. One way to resolve this surprise is to find some reason to believe these strange abilities are fundamentally not an understanding of the world. Thus stochastic parrots, this article, Yan LeCun and Chomsky, etc. Another way to resolve this surprise is to find some reason to believe these strange abilities fundamentally are an understanding of the world. Thus regulation of AI, existential risk, Hinton and Yudkowsky, etc. I don’t know what the correct resolution of the surprise is. The only thing I’m confident in is that it’s correct to be surprised by the abilities of LLMs. My current (tentative) resolution of the surprise is that language encoded way more information about reality than we thought it did. (Enough information that you can fully derive reality from language seems improbable, but iirc it did derive Othello and partly derived chess and I would have thought there wasn’t enough information in language to derive those without playing the games as well, so I can’t rule it out.)
- JoshTko 3y agoThe more I think about it the more I'm convinced I am basically just predicting/saying my next word whenever I speak.
- morkalork 3y agoWhy not? Your brain is already making predictions about what you expect to see and hear as a part of your perception of reality anyways.
- quad_eye_oh 3y ago
- cjbprime 3y ago> Brooks: No, because it doesn’t have any underlying model of the world. I don't know whether to be more disappointed with the famous technologists who are apparently unable to think of questions to ask GPT-4 that require a world model to answer, or with the writers who don't question them about it.
- pama 3y agoUnfortunately even though the questions are about GPT-4 the answers and personal experience only refer to GPT-3.5 at most. I hope openai changes the name of the next version to avoid this confusing narrative by prominent people. 3.5 vs 4 is like comparing a toddler to a high school kid.
- osigurdson 3y agoThe primary question is, what phase are we in now? Are we just at the beginning or somehow already in the asymptotic phase?
- idopmstuff 3y ago> The example I used at the time was, I think it was a Google program labeling an image of people playing Frisbee in the park. And if a person says, “Oh, that’s a person playing Frisbee in the park,” you would assume you could ask him a question, like, “Can you eat a Frisbee?” And they would know, of course not; it’s made of plastic. You’d just expect they’d have that competence. That they would know the answer to the question, “Can you play Frisbee in a snowstorm? Or, how far can a person throw a Frisbee? Can they throw it 10 miles? Can they only throw it 10 centimeters?” You’d expect all that competence from that one piece of performance: a person saying, “That’s a picture of people playing Frisbee in the park.” This seems like exactly a set of things that GPT-4 can do. The image recognition capabilities haven't been released yet, but they were demoed when it launched and clearly have the ability to handle a situation like this. From there, you could ask it every single one of these questions and get the correct answer.
- ck2 3y agoDifferent perspective from MIT https://www.technologyreview.com/2023/05/02/1072528/geoffrey-hinton-google-why-scared-ai/ https://www.technologyreview.com/2023/05/02/1072528/geoffrey...
- zone411 3y agoRodney Brooks' deep learning predictions have not aged well. For example, in his 2018 blog post (http://rodneybrooks.com/forai-steps-toward-super-intelligence-i-how-we-got-here/ http://rodneybrooks.com/forai-steps-toward-super-intelligenc...), he rates various approaches from 1-3 (with 3 being the best). Neural nets scored: Composition: 1 Grounding: 3 Spatial: 1 Sentience: 1 Ambiguity: 2 At that time, the potential of neural nets was already very clear. He also predicted that by 2020 we'll have popular press stories that the era of Deep Learning is over and that by 2021 VCs will figure out that for an investment to pay off there needs to be something more than "X + Deep Learning".
- p-e-w 3y agoI'll let you in on a secret: There aren't actually any "AI experts". There are machine learning experts, that is, people whose expertise lies in designing and analyzing systems that perform (semi-)automatic inference on data. But nobody can be an expert on "artificial intelligence", because we don't know what that word really means. We don't even know what intelligence really is. We have no idea how the human mind works. We don't understand emergence, at all, which is why we continue to be surprised when it happens. So it shouldn't come as a shock when eminent figures commonly labeled "AI experts" make predictions that turn out to be fundamentally and embarrassingly wrong in a very short timeframe: They're just talking out of their behinds, like everyone else.
- xvilka 3y ago> We have no idea how the human mind works. True, but we already know the so called "neural networks" that many computer scientists believe are how brain works aren't even close. They are all based on half-a-century old concept of neuron that was debunked many times over, experimentally by real neuroscientists.
- p-e-w 3y agoThat's correct, but it doesn't mean artificial neural networks cannot achieve intelligence, or even superintelligence. The fact that the human brain doesn't work like that doesn't automatically imply that (structurally) simpler models are fundamentally less capable.
- thatguyknows 3y agoAs mentioned, much of our discussion is rote parroting. I can usually go into any hackernews thread and roughly know what the top discussions are going to be. It's not surprising that an AI trained on a large portion of the internet would thus look human like. If you really poke at GPT, you begin to realize it's fairly shallow. Human intelligence is like a deep well or pond, where as GPT is a vast but shallow ocean. Making that ocean deeper is not a trivial problem that we can just throw more compute or data at. We've pretty much tapped out that depth with GPT4 and are going to need better designs. This could only take half a decade or it could be half a century. Plenty of enterprises stagnate for decades.
- p-e-w 3y ago> Making that ocean deeper is not a trivial problem that we can just throw more compute or data at. You can't possibly know that, given that we don't actually understand how LLMs work on a high level. > We've pretty much tapped out that depth with GPT4 GPT-4 is three months old and you're confident that its working principle cannot be extended further? Where do you get that confidence from?
- thatguyknows 3y agoSam Altman said it himself. He seems like a reasonable source. If you're familiar with other fields of AI, adding more and more layers to ResNet was the hotness for awhile, but the trick stopped working after awhile.
- steveBK123 3y agoExactly, and OpenAI has been around nearly 8 years, consumed huge amount of data with tons of compute. They are just showing us the product now. It is possible they've reached some 80/20 point and he is pretty honest about how much more extendable the current approach really is. Would explain going to congress and asking for regulation (of their not-quite-there-yet competitors who they want a regulatory moat against).
- 3y ago
- p-e-w 3y ago> It gives an answer with complete confidence, and I sort of believe it. And half the time, it’s completely wrong. That's bullshit, unless you are asking questions specifically designed to make GPT-4 hallucinate. For most real-world, everyday topics, the accuracy is close to 100%. GPT-4 would be utterly useless otherwise.
- m3kw9 3y agoYeah a lot of times is right, unless they are really complex subjects then maybe even less then 50
- munchler 3y agoWhich is true for many humans as well.
- thorum 3y agoLess about complexity than about how well-documented the subject is on the internet IMO. I’ve been using it to help me set up and troubleshoot AWS Elastic Kubernetes clusters, which are plenty complex, and I’d estimate it’s been around 95% accurate. (And for at least one of the times it seemed to be wrong, it turned out I’d made a mistake following its instructions...)
- galaxyLogic 3y agoYou could ask me a difficult scientific question which I wouldn't even understand. But I could google and find a scientific paper which I would pass to you. You could say fantastic answer, thanks. But I would have no clue as to whether it is or is not. Now if I could just do that fast enough to serve all the people all the time, you would call me a sensation. I think this is what's happening with LLMs.
- helloplanets 3y agoSuch a weird time, when the gap in the performance of GPT 3.5 and 4 is huge, but the time between their releases is so short. Some of the critique that was apt for 3.5 sounds a bit out of touch when it comes to 4.
- jmyeet 3y agoGPT-4 is pretty amazing but I, too, feel this is being overhyped. For me, a sobering example is how OpenAI does math (eg [1]). Specifically, the model clearly doesn't really understand multiplication and "learns" it from training data. This tends to get the first few and last few digits right for a simple multiplication with 6-7 digit numbers. Now you can solve that with plugins (eg training the model to recognize math problems and have access to a calculator) so it's a solvable problem but you realize there's an extremely long tail of such problems. It goes to show that GPT-4 isn't "magic" and we still have a long way to go. [1]: https://www.reddit.com/r/OpenAI/comments/12donja/gpt4_and_math/ https://www.reddit.com/r/OpenAI/comments/12donja/gpt4_and_ma...
- seanhunter 3y agoMost of the time when people find a maths problem that they can trick the model into getting wrong, it's also possible to get the model to give the correct answer with better prompting. A trick that's worth knowing is just to ask the model to give each step in the solution and explain as it goes. This gives the model "time to think" and leads to better results.
- usaar333 3y agoPretty sure you can't get GPT-4 to do 8 digit multiplication with any prompt. For what it's worth, I'm not even sure if chain of thought provides much value to GPT-4. The RLHF it went through seems to have encouraged more logical thinking already.
- morkalork 3y agoNow you have me wondering if you could prompt it to solve it step by step, using the same method an elementary student would do on paper.
- seanhunter 3y agoYou can. One thing that I've found fun is to prompt it for some maths problems without solutions, then I provide solutions and for any that I get wrong ask it to explain my mistakes.
- 2devnull 3y agoIf it weren’t for gpt all we’d be talking about is layoffs so thank you very much i’ll take ridiculous ai talk over the alternative.
- thomastjeffery 3y agoBrooks has done something I appreciate a lot: he turned a phrase. > stop confusing performance with competence You can safely skip the rest of the article. That sentence gives you all you need, because you are competent. If you want a little more meat: > The example I used at the time was, I think it was a Google program labeling an image of people playing Frisbee in the park. And if a person says, “Oh, that’s a person playing Frisbee in the park,” you would assume you could ask him a question, like, “Can you eat a Frisbee?” And they would know, of course not; it’s made of plastic. You’d just expect they’d have that competence. That they would know the answer to the question, “Can you play Frisbee in a snowstorm? Or, how far can a person throw a Frisbee? Can they throw it 10 miles? Can they only throw it 10 centimeters?” You’d expect all that competence from that one piece of performance: a person saying, “That’s a picture of people playing Frisbee in the park.” --- So I've calmed down. Now what? The problem isn't only that this train is flying off on a tangent: it's that it's off the rails. What rails should it be on? The problem, as I see it, is narrative. As soon as we called it "AI", that wrote the Genesis of the Scripture of the cult. In this new religious movement, God is spelled L-L-M. Back here in reality, LLM isn't a God; or even a person at all. That's the mistake: personification. A person can perform, but a performance can't person. --- Narrative is a powerful tool. It's why we're so excited about Natural Language Processing in the first place. Ever since the very origins of software, the power of narrative has been so close, but always still just out of grasp. Do we even know what we are reaching for in the first place? In a sense, we have a part of it: explicit definition. What Chomsky categorized "Context-Free Grammar", we have made into programming languages. What they are missing is implicit inference: context. That's what LLMs do. They use inference to model the patterns that exist in written text. With that model, they can hallucinate more text that follows the same patterns: they can perform natural language. So that's it, right? Problem solved! What's missing? explicit definition. We traded one problem for another. No one (so far) has figured out how to solve both in the same program. You can have definition, or, you can have inference. You can't have both. This doesn't make any sense to us humans. We don't have any trouble at all doing both at the same time. We do it all the time! Do we actually do anything else? Unfortunately, LLMs are not humans. --- The two approaches to language are diametrically opposed, but they work with the same domain. Approaching from either end of the spectrum, definition and inference explore together the wild universe that is story. That's the missing piece: once we figure out what story is made of, we should be able to put all three pieces together.
- mitthrowaway2 3y agoNot really the thrust of the article, but that '50s picture of the family playing scrabble in a self-driving car, surrounded by text about trains... really makes me think that if I were that family, I'd still prefer to be playing scrabble on a train, rather than in a cramped self-driving car on a highway.
- m3kw9 3y agoGpt is a great tool, it won’t be able to do complex tasks because it really isn’t that smart. If you tell it to do something relatively complex end to end, it will fail unless a plugin specifically supports it.
- ijidak 3y agoThese debates about how well GPT can think seem merely philosophical. This is a perfect case of perfection being the enemy of the good. Useful AI is here. Hard stop. The impacts will be huge and unpredictable. Billions will be made. The world will change. Humans will continue making rapid progress, via merging various AI methods and new breakthroughs. Nevertheless, I enjoy reading the debate. But anyone wringing their hands over how much is GPT thinking is missing the point. These are companies making products. This is not academic research. It's just another tool in a long list of tools made by humans. And it's already a productive tool. This reminds me of the skepticism surrounding electric cars while Tesla was already growing by leaps and bounds. The ship has sailed. The revolution has started. Progress will undoubtedly be rapid and continual.
- namaria 3y ago> Nevertheless, I enjoy reading the debate. > But anyone wringing their hands over how much is GPT thinking is missing the point. You seem to be missing the point. > These debates about how well GPT can think seem merely philosophical. Merely? Yeah, you're missing the point. You want a debate or you think the debate is meaningless? You don't get to appreciate it and call it pointless and sound reasonable at the same time. > The ship has sailed. The revolution has started. Progress will undoubtedly be rapid and continual. It started 2 million years ago when humans started roaming the planet. We're clearly a runway process. We don't need a chat bot to prove it.
- jiggawatts 3y agoThis is a terrible article written by someone who doesn't seem to have even tried GPT 4. Their only example references GPT 3.5, for example, and then they waffle on about only vaguely related topics such as level 5 self-driving. This quote in particular stood out as ignorant: “What the large language models are good at is saying what an answer should sound like, which is different from what an answer should be.” That's... not at all how large language models work. Tiny, trivial, toy language models work like this, because they don't have the internal capacity to do anything else. They just don't have enough parameters. Stephen Wolfram explained it best: After a point, the only way to get better at modelling the statistics of language is to go to the level "above" grammar and start modelling common sense facts about the world. The larger the model, the higher the level of abstraction it can reach to improve its predictions. His example was this sentence: "The elephant flew to the Moon." That is a syntactically and grammatically correct sentence. A toy LLM, or older NLP algorithms will mark that as "valid" and happily match it, predict it, or whatever. But elephants don't fly to the Moon, not because the sentence is invalid, but because they can't fly, the Moon has never been visited by any animal, and even humans can't reach it (at the moment). To predict that this sentence is unlikely, the model has to encode all of that knowledge about the world. Go ask GPT 4 -- not 3.5 -- what it thinks about elephants flying to the moon. Then, and only them go write a snarky IEEE article.
- zmnd 3y agoI think the main reason for division is that everyone projects to their own use cases. I have been using gpt-4 for quite some time and also couldn't understand why someone would say that it just produces something that sounds like a real answer. But then I found some queries that can definitely be described as "sounding like truth". So your personal experience probably wasn't what was their experience. For those curious, I was asking gpt-4 about the top 3 cards from my favorite board game, Spirit Island. All three of them sounded really convincing, having the same structure and the same writing style, but unfortunately none of them existed. So everything that fails outside of most common use cases would probably have an experience of convincing hallucinations.
- jiggawatts 3y ago
- osigurdson 3y agoI don’t completely disagree but this article has a lot of “they said x about y, and that didn’t come true” type arguments which don’t resonate.
- galaxyLogic 3y ago> it answers with such confidence any question I ask. It gives an answer with complete confidence, and I sort of believe it. And half the time, it’s completely wrong. That's it. LLMs don't have SHAME! They simply don't care if what they're saying is false or not. They are like some politicians of late. They don't even understand that giving out a wrong or misleading answer will affect their credibility. You see LLMs don't have a DESIRE for credibility. They don't have desires. We need not fear these models. But we do need to fear some people who will use them for evil purposes.
- Animats 3y ago"And I think what they say, interestingly, is how much of our language is very much rote." I've been saying that for a while. Large language model systems have made it painfully clear that much of what humans thought was intelligent behavior is rather banal. The scary thing is that a sizable fraction of white-collar work is banal enough to be done by such systems.
- gmerc 3y agoieee has been on a copium suspension trip for a while now.
- braindead_in 3y ago> No, because it doesn’t have any underlying model of the world. Ilya's counter to this reasoning is for next word prediction to work, the model has to 'understand' our world. Otherwise the predictions will be way off. Therefore the human world has been modelled to a degree by GPT. I haven't yet heard a good counter to that.
- arisAlexis 3y agoAmazingly wrong and proven wrong. This guy didn't get the memo where GPT passed all the medical and law exams and now coding. Coding is not "how an answer should be like", it is the answer. When humanity will stop hearing out peripherals to AI? Let's focus in what Ilya says or Hinton.
- stareatgoats 3y agoOne of my goals in life before dementia sets in for real is to devise some model, perhaps a conceptual framework which will allow us to escape the clutches of habitual simplification, a subset of which is dichotomous thinking, which in turn leads to the inevitable painting of strawmen as a way to prove our point (among other things). How sweet it would be to shortcut all the mandatory twists and turns of discourse that follows: "this is a mischaracterization of x", "not all x are y", "x and z are really not opposites, but overlapping", "this is a spectrum with a bell curve, not an either/or" etc. But of course, we all do this, not because we can't think clearly, but because we have an agenda, or maybe more frequently: want to trash talk the stance of an opponent because of what the proliferation of that stance might lead to, and so on. Taking such into account should be an integral part of the conceptual framework, obviously. In the case at hand, one could easily argue that people in the debate are creating false dichotomies: LLMs are either stochastic parrots OR algorithms with an understanding, when in reality they are both (and also something else completely), but acknowledging such would likely require that one doesn't have an axe to grind, a stake in the field or what you might call it. It would require extending some "philosophers charity" to an opponent, that maybe has tried to undercut one's work for decades, in a field steeped in fierce and bitter competition for a name, like academia. Or, in case one has a business in the field, it would require maybe saying something that puts your core business idea in the crosshairs of legislators, or something else that doesn't serve your long term business interests. Which brings us to this important aspect of this "conceptual framework against simplification" already briefly touched upon, namely identifying the bias of the participants in the debate. My impression is that naming bias has largely gone out of fashion, which is a pity because it is really a necessary part of understanding an argument: it rarely explains it all (that would be a grave simplification), but it is really a vital part of understanding an argument. And a difficult one, because people will go to extreme lengths to hide their agenda. And the current conceptual framework for unravelling bias has largely been occupied by the fact-checking industry: i.e. things are either true or false, and once you are cleared (like most mainstream media) then bias is not questioned. But we can be assured, there is always some bias, and it is usually relevant to name it (if one can see it), even if it infuriates the named party. Just sayin'.
- namaria 3y ago
- deleted 3y ago[deleted]
- bohadi 3y agoapropos the roomba founder, a nontechnical argument for necessity of embodied AI circa 1980 (5 min video) https://youtu.be/QMMw9fQ452c?t=49 https://youtu.be/QMMw9fQ452c?t=49 That LLMs learn a world model is very convincing now, but as LeCun has said it's just one piece of the intelligence puzzle, incl. perceiving, actuating shenmede