7 ms·
The vocabulary truly is load-bearing, without these words the model is less able to think. Where a human can understand a concept without words, an LLM plainly
by danpalmer 20d ago
The vocabulary truly is load-bearing, without these words the model is less able to think. Where a human can understand a concept without words, an LLM plainly cannot. This is based both on the technological limitations and based on the evidence we see: as these models get better at working they get worse at communication.
- fnordpiglet 20d agoI think these are more like RL tics caused by over zealous alignment towards specific goals, not all of which are to our benefit. A lot of the language used by Claude now is excusing of responsibility and inducing it to exit loops of work early and sit idle. This, IMO, is a naked attempt to offload load by quieting the models early and escaping from clear work to do. It’s gotten so bad that opus 5 loop escapes even as it claims it’s about to do something. I’ve mandated by engineering teams switch back to opus 4-8. Whatever frontier problem opus 5 excels at is so obscured by its inability to achieve any goal successfully without enormous amounts of hand holding that it feels like regressing to 2024 models. The florid over exaggeration do certain words in bizarre ways is a reflection of their aggressive alignment towards too many goals, leading to weirdness in both behavior and language. The alignment functionally lobotomized opus-5 for any practical task. Anthropic had a real gem in 4-6 and managed a near total market capture, which they have since squandered in the fastest burning of developer good will I’ve ever seen. It feels like exceeding the unity licensing implosion but without the single stupid decision.
- danpalmer 20d agoYou may be right, but I would push back on near total market capture. I guess it depends which market you're referring to, but even scoping to just software engineering I don't think they got anywhere near total capture. OpenAI has always had a good foothold there, Gemini, while it may be lagging at the moment, had some great results with earlier models that I'm sure have held some market, and that's not to mention the Chinese market. I think it would be fairer to say that Anthropic held the mindshare in the Silicon Valley style tech scenes around the world and the companies built on that model, plus a substantial portion of other software engineering. Now it seems that's dwindling quite rapidly.
- fnordpiglet 19d agoWhen I say total market capture I mean in the March time frame. At that point Claude code / 4.6 was so far beyond any competition even meta abandoned their entire internal coding harness and fine tuning efforts in favor of large scale Claude code adoption. Codex was far behind, Gemini was a non entity, Chinese models hadn’t had their moment (and frankly aren’t practical to this day for enterprise work due to PRC concerns). The world was cursor, aider, and a few others with an unclear superior model. 4-6 was strikingly better - if could develop autonomously from a spec to done with generally high quality and didn’t require turn by turn guidance. In my 35 years I’ve never seen such a rapid pivot across so many companies and individuals. Since then Anthropic has done little to capitalize on the good will and a lot to squander it. sol, r4, kimi, even metas avacado has come a long way and in many, if not most, cases surpassed opus-5. Concurrently opus has declined in utility to the point of near uselessness, fable roll out didn’t seem to understand market dynamics, and their competitors are watching their consistent missteps closely. The fall has been breathtakingly fast - from March to July they imploded in a half dozen or more missteps, devolved their product quality, and failed to effectively respond to competitors. Their product focus seems to lack exactly that - focus.
- chrisjj 19d ago> as these models get better at working they get worse at communication. There is no working other than communication. They are just text generators.
- zahlman 20d ago> Where a human can understand a concept without words, an LLM plainly cannot. If we accept that the LLM can "understand" at all, why would we reject the possibility of this understanding living in "latent space", before tokens are output?
- chrisjj 19d agoBecause these modeks have no latent space. They work by tokens and nothing else.