5 ms·
"Next-token predictor" is one of those phrases used most of the time with a motive to downplay the abilities and faculties of AI models. It is intended to trivi
by atleastoptimal 12d ago
"Next-token predictor" is one of those phrases used most of the time with a motive to downplay the abilities and faculties of AI models. It is intended to trivialize LLM's and imply that there is some fundamental limit on their capacities.
Relying on it as a mental model for what LLM's are minimizes the emergent properties of scaling. It's like imagining that unicellular life could never eventually evolve into complex multi-cellular organisms because individual cells are just "survival and next-mitosis optimizers"
- jvanderbot 12d agoBut it is a next token predictor. Recursively invoked. With carefully selected context. And massive investment in RL to tune token selection. And the ability to use cli tools on other folks' machines. That's a powerful system built around a conceptually simple technology: Next token predictors.
- atleastoptimal 12d agoYes this is correct. The thing is not about the term next-token predictor being correct, but because of the connotative weight of that phrase as a implicit trivialization of LLM abilities, which is how it is often used.
- noduerme 12d agoWhat is the motivation behind advocating against people trivializing LLMs? As in, why do you care?
- mofeien 12d agoNot the parent, but this incorrect trivialization of LLMs is often employed as a counterargument to the risks of AI such as "will take your job" or "will escape human control (again and worse)" or just "can possibly hurt me". And taking the easy feel-good cop-out instead of actively engaging with these questions is just.. harmful?
- noduerme 12d agoAh. I always thought it came from the LLM booster perspective of trying to prove emergent intelligence. A lot of very clever autocompletes working together can be incredibly dangerous.
- 27183 12d agoFrom another point of view, campaigning against the "next token predictor model" is a means to implicitly inflate LLMs' abilities. Given all the other hype-inducing terminology we've seen--"reasoning", most egregiously IMO--this seems more likely. Is there a simple, more accurate mental model? From what I've seen of the literature, "next token predictor" is a very accurate first order description of what an LLM does, I can't really do better, therefore this or that connotative interpretation isn't giving me a great deal of pause.
- mort96 12d agoAt the same time, it ... is literally a next token predictor. Like that's what it is. The input is a sequence of tokens. The output is a probability distribution of next tokens.
- doc_ick 12d ago100%
- gjm11 12d agoIt is. And human beings are bags of chemicals. But for many purposes you will not find it helpful to think of human beings as bags of chemicals, and for many purposes you will not find it helpful to think of LLMs as next-token predictors.
- Aurornis 12d ago> But for many purposes you will not find it helpful to think of human beings as bags of chemicals But when we talk about humans, we're not talking about the chemicals involved in those humans. When we talk about LLMs, the tokens are the valuable thing they produce for us. We want LLMs because they give us sequences of tokens.
- jayd16 12d agoIt can be pretty helpful to think of human function in chemical terms. Its at least unhelpful to deny it.
- weego 12d agoimply that there is some fundamental limit on their capacities This is a wildly dismissive statement that does a lot of heavy lifting. Your assertion is that we just happened to hit on a methodology that has no limitations between being an encyclopedia with a novel human language interface and, I guess by implication, AGI? That seems more outrageous a claim than the one you're dismissing.
- atleastoptimal 12d agoI don't think it's outrageous when many of the people who claimed it was a next-token predictor have been proven wrong repeatedly over the past 5 years. There were people years ago who claims AI could never answer questions like "what would happen to a ball on a table if I moved the table" correctly because its text-base world model could never intuit physics, or that it could never do math or code accurately. When I say there is some issue with people claiming there is some fundamental limit on the capacities of LLM's, I don't mean to say "If you think that they don't have unlimited potential you are wrong", I mean "you can't use the architecture of the transformer to make a sweeping declaration of things LLM's can or cannot do without empirical evidence, because the empirical evidence has unearthed far more surprising revelations than a reductive theory has been able to"
- uludag 12d agoI'm actually extremely confident that I can use the architecture to make a sweeping claim on what it can or can't do and will be extremely surprised if proven wrong: A pure next-token language model won't be able to give detailed instructions to an ensemble of motors, mimicking a human body, to do a wide variety of tasks our human brain is excellent at doing, for example, inserting keys into a car, opening the door, sitting down, starting the car, putting the car in reverse, and exit a parking lot, being careful not to hit anything.
- bigstrat2003 12d ago> it could never do math or code accurately. They still can't do code accurately. The fact that you use this as a defense of your position greatly undermines the credibility of your claim.
- pjerem 12d agoGood example. It’s also like saying our brains are just electric circuitry incorporated in meat. It’s true but it seems that consciousness emerges from this. The fact that LLMs are next token predictors isn’t the interesting or impressive part. Actually my brain strictly is a black box predicting (or choosing) my next word/action/move… based on a complex existing context (my thoughts, the environment, my physical state, my senses…). FWIW, I don’t believe LLMs are sentient, but I don’t think either that we have enough knowledge to rule it out.
- mmoll 12d agoThat is the point: our minds are also next-„token“-predictors, at least we can‘t prove they‘re not. That‘s why I don‘t agree with the article: LLMs _are_ next-token predictors. However, that says little about their capabilities. Also, while I have no idea what „consciousness“ is, I have difficulties believing that it could arise in a program that, in theory, you could execute with pen and paper.
- otabdeveloper4 12d agoWe also can't prove that our minds aren't machine elf meat puppets. Come on. Please.
- Dylan16807 12d ago> our minds are also next-„token“-predictors, at least we can‘t prove they‘re not Your mind can pick a random number without outputting it, participate in a short conversation, and then say the number.
- otabdeveloper4 12d ago> It’s true It's not. "Brains as electrical circuits" is a gross simplification based on our ignorance and prejudices. (In the 18th century they spoke of brains as "clockwork mechanisms".) LLMs, in contrast, are literally next token predictors. We know exactly how LLMs work, and they are exactly that.
- 12d ago
- otabdeveloper4 12d agoThat's literally what LLMs are. No amount of cope and anthropomorphizing is gonna change that cold, hard fact. P.S. The perceived magic of LLMs comes from the way they cross-correlate all the probabilities of tokens on their context window. Not from their ability to "think ahead". They can't do that by design.
- uludag 12d agoAnd what's wrong with downplaying the abilities and faculties of AI models if that's what people feel like saying? We don't call humans or animals sacks of chemicals because we believe they have moral status.
- gruntled-worker 12d ago> used most of the time with a motive to downplay the abilities and faculties of AI models Exactly. We're dancing around the real argument: there's massive amounts of influencing going on (and not only about AI.)