5 ms·
ITT still, among the highly technical, a surprising incomprehension of the fact that describing LLM in terms of "token prediction" and dismissing this as statis
by aaroninsf 26d ago
ITT still, among the highly technical, a surprising incomprehension of the fact that describing LLM in terms of "token prediction" and dismissing this as statistical,
neatly dodges literally everything that is interesting about what they do.
You to, friend, consume inputs and generate outputs.
What's interesting is how you do that, and, if you prefer to look at it through a technician's lens, whether or not what is done is reducible.
I don't mean quantizing the model, great, now you have a crankshaft with no oil. But the motion of the pistons and wheels is roughly the same.
What I mean is, to put it in plain terms, the only way you find out what a model is going to "predict" from a given input is to ask it.
If your mental model is still that LLM are "glorified overhyped giant markov chains" performing "parroting" you need to improve your understanding.
- Anamon 24d agoCan you explain what you think makes today's LLMs more than glorified Markov chains? I know they're not technically Markov chains, but conceptually, it's still close enough. There's nothing revolutionarily new. No amount of loops and harness tricks is going to change that fact.