7 ms·
> Cracks are starting to appear. Open Ai and Anthropic are publicly asking for slowdown in AI research. Translation: We see this technology not being any more u
by dahart 4d ago
> Cracks are starting to appear. Open Ai and Anthropic are publicly asking for slowdown in AI research. Translation: We see this technology not being any more useful than what it is now, no AGI is coming
Predicting a usefulness plateau is absolutely wild given how fast AI agents have been improving at writing code this year. I have doubts about AGI but I think you’re making assumptions and translating it wrong. These two companies have always been asking for a slowdown from their inception, that’s not a new thing. It’s part marketing hype, but they both do want regulation to step in and slow down the competition, not because they see a usefulness plateau, but the opposite - the usefulness is growing so fast that they want to remain in control, and they are scared that working hard and competing will not be enough. Anthropic has also said out loud they think their competition (not just OpenAI) is not being responsible and they want the regulation so they can be the responsible shepherd, as AI gets more and more useful.
- tripledry 4d agoI agree with you generally, just an observation on coding specifically. Have the models improved since Opus 4.x? I find the newer models are not better in my day job, maybe in one shotting mvp's and other tasks. Not trying to argue your point, just intrested in the coding aspect, if the models were improving as fast as benchmarks I would expect capability improvements to be obvious, but talking to people and reading forums, it seems everyone has a different opinion.
- blfr 4d agoYes, benchmarks are gamed and only loosely indicative of real world performance. Also yes, Fable is massively better than Opus. It requires significantly less instruction and specs and produces more directly mergeable code. The improvement is obvious as soon as my Fable allotment runs out and I try to do something with Opus. Have you given the same (larger) task to Opus and Fable?
- tripledry 3d agoI have not experimented that much, my observation is mostly that people seem to have different experiences with the capabilities. Personally I still use mainly Opus and find it handles most tasks quite well (without burning all my corporate quota).
- Culonavirus 4d agoAstra, on average, despite its GPT 6 version bump, is not any better at coding than Sol. Some even argue its worse in practice due to the varying quality of its output. https://x.com/theo/status/2097192907023458473 https://x.com/theo/status/2097192907023458473 (^ this "swearing at a model" thing has happened to me multiple times on Astra already) If you have to "debate" the quality of a new Big Number model (and double and triple check your eyes and model setting switches when it pukes up complete garbage), that is NOT a good sign.
- heaney-555 3d ago>Astra, on average, despite its GPT 6 version bump, is not any better at coding than Sol. This is a ridiculous claim that is disproven by simply using it for more than 5 minutes. https://withspecific.com/benchmarks/real-swe https://withspecific.com/benchmarks/real-swe
- deleted 3d ago[deleted]
- anu7df 3d agoFor me opus-4.8 was the most useful coding model. Sure fable is better at planning but is expensive to use as a daily driver. But even fable, when it fails, fails in such strange ways that I am now convinced this intelligence is an illusion and path to AGI lies elsewhere. I was willing to buy the whole emergent intelligence claim till last year. Now, not so much. As for calling for slow down, sure the two companies kept parroting each other's lines but they were not slowing down the cash burn or gpu purchases. Why would they? The prize was too high. Now, with open Ai needing trillion dollar valuation to IPO and private funding possibly showing signs of slowing down (only for these labs because the valuation is too high to begin with for most prudent investors: My speculation) they have no option but to slow down. Then would you rather say, slowed down because we are running out of cash or that we are slowing down because national security? It really cannot be that they can't solve alignment but otherwise it is really powerful and improving. Simply because airgap exists and we know how to do it. If all else fails power off the freaking gigawatt cluster. More likely this recursive self improvement is an unstable loop and the model is likely degrading with self improvement effort. At 100's of millions per experiment this is going to be unsustainable.
- dahart 3d agoAh, now having the best model be expensive to use, and thinking the cash burn of both companies is insane and unsustainable I completely agree with. This, I think is likely the biggest reason they’re asking for government intervention and regulation: to help them weather the coming investment/cash plateau, not because the models are nearing any asymptotic limits. To be fair, there is an argument to be made that the improvement in models is tied to the cash burn; if they can’t train bigger models or acquire more data or research more effective harnesses, then the improvement of their models might slow down - while GLM or other models continue to improve. I’ve never thought LLMs were on the AGI path, but I have to admit it’s surprising how far it’s come with no end in sight yet. There is something important to be said about how ‘intelligence’ is embedded in language, and it suggests that intelligence isn’t exactly what we thought it was. The language component of intelligence also goes a long way to explaining technology’s progress in human civilization; how language and the printing press and mail and radio/tv and the internet have each ushered in accelerations in the pace of progress. Biologically and evolutionarily speaking, it’s unlikely that humans have become any smarter in the last two thousand years, but technology (among other things) has exploded.