9 ms·
They are talking about slowing down the public facing AI development. Because then nation states can create a capabilities gap between them and the public. Why
by AaronAPU 4d ago
They are talking about slowing down the public facing AI development. Because then nation states can create a capabilities gap between them and the public.
Why does nobody seem to be pointing out this obvious explanation? It explains why the “we need to race China” concern suddenly vanished in the discussion.
The government can simply gag Sam, Dario, Musk on national security basis, getting them all behind the public messaging.
- bossyTeacher 3d ago> They are talking about slowing down the public facing AI development. I assume there already existed advanced models frontier labs only made available for national security. We just weren't told. I am guessing the announcement is a way to avoid jeopardising their IPO, while reducing their workload.
- ranger_danger 4d agoI figured they're just admitting AI models have plateaued and are coming up with some fake story about self restraint so they don't lose VC money
- lukan 4d agoNot sure about the level of irony here, but I keep hearing models have plateaued since a while now, but I keep being impressed with the latest model performance.
- unleashhale 4d agoAnything in particular? My experience has been like seeing the addition of retractable cupholders, but maybe different domains.
- qskousen 4d agoI have a pet project I have been working away on for some time that involves building GPU backends for various cards in Zig, lots of complex stuff in it. Lately I mostly use Opus 5, it can pretty reliably plug away at things but it does mess stuff up occasionally. For this codebase, Fable 5.1 was noticeably better at getting things right and doing things in a good reliable way. Of course, I can only use Fable for a bit before I hit the usage cap for the week, so I save it for the tougher things. That said, I absolutely abhor the way recent Anthropic models write prose, especially comments. I recently tried doing a fairly normal task for this codebase with codex, as I have seen a lot of people talking it up on here. A single task running for ~1-2 hours burned through over half of my usage for the week on the $125/month plan, not on a top model (I don't remember which one specifically I used). It struggled to get the basics done, then got absolutely stuck on a follow up. Handed it over to Claude and it 1-shot it.
- eru 4d agoI really liked codex in the last few weeks, especially its ability to clean up after Claude's (prose) messes and do reviews. But in the last few days something seems to have happened that made Codex's models massively stupider (for what I am doing). Really weirdly, it suddenly refused to even run tests it previously wrote itself (and previously ran), because of some false positive about cybersecurity. That by itself is not evidence of stupidity. Trying to make a 200+ file PR full of research notes is, and the PR didn't even solve the problem I asked it to.
- ipaddr 4d agoThat's the case for open weights models. Hosted Deepseek Flash 731 copy isn't changing randomly one day because the parent company decided to change it.
- drTobiasFunke 4d agoReally? My employer rolled back to opus 4.8 because 5 was expensive AND crap. Didnt even consider fable because it didn’t add any additional value. For most software eng and design work opus 4.6-4.8 just works fine. For everyday joe asking ai to plan a trip or home diy work even sonnet works fine. Any cybersecurity or other areas are niches that cannot support trillion $ valuations. What am I missing? Genuinely curious
- bpodgursky 4d agoIf Fable doesn't add additional value in your workplace, it means you aren't being ambitious enough in how you integrate agents into your workstream. Yes, it's probably comparable to 4.8 if you are just using it to write code and put up a couple pull requests. That's not where things are now.
- bix6 4d agoAnd where are things now?
- nradov 4d agoYou shouldn't be down voted, AI native companies have already moved up to the next level beyond writing individual PRs.
- Tanjreeve 4d ago"AI Native" here meaning Companies where your token use isn't scrutinised/capped yet?
- nozzlegear 4d agoThis is just "you're holding it wrong" with a little smooch of condescension. If only we plebeians could comprehend what magnificent works those who have ambitiously integrated agents into the workstream have wrought!
- bpodgursky 4d ago
- 0xcde4c3db 4d agoI don't think "plateaued" is the right word, but I do feel like there's been something like a logistic curve compression in the difference between smaller and larger models as the field evolves. For inference at least, the scale of practical difference between a single high-VRAM GPU or SFF UMA box, a whole rack, and a whole data center seems to be falling far short of what we might have imagined just a few years ago. The conversations I've heard have largely turned away from breathless anticipation of the next frontier model and toward attempts at hard-nosed evaluation of which tokens are worth the cost.
- eru 4d agoMaybe, though that's another kind of progress in itself. Very impressive progress!
- antupis 4d agoI think it’s more that pushing frontier is extremely costly and there is no free lunches in same way as 2024.
- int_19h 4d agoThe difference is still vast, it's just that the smaller stuff is "good enough" for many things now.
- whateveracct 4d agoastra is more parlor tricks than real gains tbh i swear they trained in on threejs in particular so those idiots on twitter could spam their garbage demos
- sampullman 4d agoI'm sure that's part of it, but I run it side by side in my review bot, and Astra medium effort consistently catches more issues than Sol 5.6, using fewer tokens. For coding it's a little harder to tell, but at least the prose feels a little better.
- int_19h 4d agoI strongly disagree. I'm working on a semantic model for Lojban, heavily AI assisted, using multiple models. I have basically all popular frontier models doing research and panel debates. Astra and Fable are both noticeably ahead of everything else including their previous iterations. When it comes to reviews, they can also find more issues in others (or even their own) code.
- f4dd 4d agoThey've not plataued but they're certainly not as impressive as the hype would have them to be. The reality is, it doesnt matter if LLMs keep getting more powerful because they still need a human to steer it. Without the human providing inputs to the LLM it just sits there and does nothing.
- arcanemachiner 4d agoYou don't need human input. Any coherent input will do the trick. You can, for example, hook it up to a logging system and have it fix errors as they occur on your platform.
- paulhebert 4d agoHave you tried this? How did it go? I’d be curious about: - your setup. How it all works - The types of errors it fixed and how quickly - Any regressions or issues it caused - The cost Thanks!
- int_19h 4d agoI have something like that running locally for my agentic harness (which includes cross-model messaging). There's a dedicated "product manager" session for it, and all other PMs are instructed to report issues with the harness as they occur to that session, while it is tasked to automatically prioritize and address them and coordinate fix deployment with other running sessions. It works surprisingly well. The errors fixed are both genuine errors in the harness itself, but increasingly so upstream bugs (in the underlying agent apps like Codex, or in Herdr, which is used to expose uniform programmatic access to all those different apps) for which it needs to come up with workarounds. No regressions so far. The cost is hard to judge on a subscription, especially when you're running really heavy tasks otherwise that dwarf any harness work.
- dismalaf 4d agoImpressed with the model performance or the chatbot/agent performance?
- rtpg 4d agoI'll take the opposite here. If someone put in frontier AI models from like .... last june I guess? in a box and let me run it with "decent" token throughput I would be happy. I think it's worth acknowledging that the power of LLMs at this point is not really so much in the smarts, but in the coordination and the surrounding harness tech. "Written english" turning into sequences of commands[0]. The whole agentic "stuff" in general. Tools + coordination is the superpower. The reasoning... it doesn't have to be _that_ good for the rest of the stuff to work. On good codebases and infra, at least. And I say this as someone who really would rather most of this stuff disappear! [0]: programming is obviously text to commands, but there's a loooooooot of futziness that LLM reasoning has let us remove in some flows
- anon373839 4d ago> If someone put in frontier AI models from like .... last june I guess? in a box and let me run it with "decent" token throughput I would be happy. You can have that! Qwen 3.8 Flash-Next is ~Opus 4.6 and runs nicely on a DGX Spark. And that’s just an architecture preview. The Qwen 4 family is expected to arrive this fall.
- rtpg 4d agoDGX Spark is a biiiiit costly but neat to hear! Do you know what kinda throughput you’re getting on that kinda setup? (I have a secondary problem of being “locked into” Claude Code by it being good enough for me, I’d probably need to investigate the other harnesses… my impression is other harnesses are a bit more aggressively OK with nuking your setup from orbit)
- anon373839 4d agoIt is costly, especially right now. I don’t think you can make a case for it on cost savings! The throughput in a single stream is about 50 tokens/sec (a bit less for prose, a bit more for code due to speculative draft acceptance rates) and about 2,000 tokens/sec for prefill. Both numbers are flat and stable as context accumulates. That’s what finally tilted me away from the Mac Studio despite its much superior memory bandwidth. I think these numbers may improve because the model is pretty new and optimizations aren’t done.
- octoberfranklin 4d agoPeople are definitely finding new things to use the models for, and orchestrating increasingly large swarms of agents in useful ways -- every single day, especially the last few months. But the basic single-NN frontier capability has been pretty stationary since Opus 4.8. Kimi K3 is almost as good as that with open weights, which has the frontier labs terrified. The only big thing on the horizon is if we can get diffusion models working reliably; that would be a big step forward. Inception's Mercury is AFAICT the leader here. It's stupifyingly fast but has obedience/hallucination problems that the autoregressives solved ~2 years ago. So it's not ready yet but improving. Also, FFS why is Grok the only model that knows how to do parallel tool calls? Such a useful ability and nobody else trains it in. Or if they do it just doesn't work.
- jandrese 4d agoI think we are in the second knee of the S curve, simply because we are hitting the point where is not enough hardware in the world to throw at this problem. These AI companies have bought everything they can and yet the models keep growing.
- nullbio 4d agoThe models have not plateaued, and they are not even mildly close to any sort of ceiling. Right now the barrier is data and compute. Quality data can be created synthetically at an exponential rate as models improve. Humans are actively feeding them with private IP. Compute advancements will begin to skyrocket as we unlock photonic computing and materials science advancements and scale up chip fabs. This is also compounding because the AI is accelerating the pace of research, testing, development, manufacturing, etc. It's a big self-accelerating feedback loop. There is no plateau.
- konmok 4d ago> Quality data can be created synthetically at an exponential rate as models improve No it can't? Every time the labs try this we see model collapse, e.g. shoving goblins into every conversation. And I have seen zero evidence that AI is accelerating materials science in any meaningful way, let alone photonic computing.
- f4dd 4d agoThere's a lot of deluland posts about.
- nullbio 4d agoYou're not well informed. Helps to keep an open mind if you want to keep up to date.
- nullbio 4d ago> Every time the labs try this we see model collapse The latest studies demonstrate model collapse is not a given and synthetic data can be used just fine. The latest models are proof of that, they're all trained on large swathes of synthetic data. It can't be used as the -only- data source of course, but that's not how it is being used. This is an obvious conclusion, too, because there's no difference between synthetic data and the data people can create, the difference is whether that data is revealing new information about the thing the model is trying to learn. If the synthetic data is just teaching the model the same thing over and over again it results in overfitting, so it needs to be done intelligently. For example, if I have an example of a puzzle, I can generalize that example and create thousands of synthetic data examples, with different rotations/perspectives, rather than having to find the data naturally. It's not that the models are just generating data out of thin air, they're generating the synthetic data on top of real world data. The smarter the models get, the better they are at generating quality synthetic variations and finding valid synthetic variations. > And I have seen zero evidence that AI is accelerating materials science in any meaningful way, let alone photonic computing. It is accelerating how quickly researchers and engineers can do their jobs. https://news.mit.edu/2026/ai-helps-design-new-materials-that-work-in-real-world-0826 https://news.mit.edu/2026/ai-helps-design-new-materials-that... This is only the beginning, too... Look ahead a year or two.
- lstodd 4d agoidk why you think "nation states" are any better at corporate governance than poster examples of bad like f/ex Boeing. Or Facebook. Or Microsoft. Or Enron for that matter. I assure you, in "nation states", that is in gov agencies it's an order or two of magnitude worse.
- nullbio 4d ago* Frontier models need infinite high quality private IP to keep them fed. Forcing an IP theft funnel ensures big lab survival and model intelligence growth. * Open-weight models are 1month behind frontier models. Cheaper, faster, private (no IP theft), steerable (you can security harden your own software without safeguard triggers). No sane business would keep using these API services if they didn't have to. The labs stand to lose a fortune. * Dario has stacked the deck at METR, who are funded by all the same NGOs who are funded by Anthropic and its investors. METR is full of ex-Anthropic employees with massive equity stakes. If they manage to position METR as the "independent evaluator" for the industry, they control what gets evaluated, how, and who passes. * Creating a gap between what the public knows exists (model capabilities) and what is used in secret allows it to be weaponized against other nations and the public. * No requirement for public disclosure on model capabilities allows them to feign they've hit intelligence ceilings while they secretly RSI to the moon with better and better chips. * Slowly but surely, this will allow the big labs to swallow the entire economy and every single business on Earth, by cloning and automating. This, and many more reasons. The labs need to feel more pressure to be held accountable for the incidents they cause (HF incident, etc), so they have an incentive to ensure it does not happen again.
- bbor 4d agoOk but the first point is just not based in reality whatsoever, sorry.
- Turn_Trout 4d agoDario's post [1] commits to direct evaluators that can, among other abilities, expose secret RSI. He wants that made law. Do you have a source on METR employees retaining massive equity stakes? [1] https://darioamodei.com/post/we-must-pace-the-frontier https://darioamodei.com/post/we-must-pace-the-frontier
- nullbio 4d agoJoe Benton left Anthropic a day before Dario's post, to work for METR evaluations. He was with Anthropic for over a year. He did the same thing that Jacob did (big song and dance about AI apocalypse, media interviews all over the place). He managed the Scalable Oversight team at Anthropic and was the research lead for the Anthropic Fellows Program. So he has equity, and likely lots of it. Then you have Josh Engels quitting DeepMind to work for METR the day before as well, doing the exact same thing. Again, doomer drama all over socials, interviews, and so on. Did I mention METR is founded by an ex-OpenAI researcher? Now you have Demis Hassabis, Sam Altman and Dario, all circlejerking eachother on X saying "we all agree with Dario" - while they ask to be "regulated" by the company that has all of their combined equity-holding ex-employees in it. METR's salaries are listing around 500k/yr. Gee, I wonder where this non-profit with ~35 people is getting all of its money? So the fact that Dario tries to frame it as an "independent third party" is all the evidence you need to know that Dario is a pathological liar and always will be. --- Some more info: Dario's sister, president of Anthropic, is married to the co-founder of Open Philanthropy. The two largest AI doomer NGOs, Center for AI Safety (CAIS) and the Future of Life Institute (FLI), have both received many millions of dollars from them. Ajeya Cotra worked at Open Philanthropy/Coefficient Giving for roughly nine years, including leading its technical AI-safety program in 2024 and contributing to AI-giving strategy in 2025. She subsequently left Coefficient and joined METR, where she is now technical staff. Ajeya is married to Paul Christiano, who founded Alignment Research Center (ARC). Alignment Research Center donated ~$4.5mil to METR. Good Ventures is a funding partner of Open Philanthropy, who funded Jacob Coxon (the first of the Anthropic employees going viral in the media) via a scholarship.
- bpodgursky 4d ago> It explains why the “we need to race China” concern suddenly vanished in the discussion Do you think you can just manifest narratives into existence? Like half of Dario's letter, that kicked off the whole thing today, is about China and how to either beat or coordinate with China.
- reverius42 4d agoSeems obvious to me, though, that you can't beat China by pausing when they don't, and China is not likely to coordinate.
- nradov 4d agoChina might pretend to coordinate in public but you can assume their secret research labs are racing full speed ahead.
- reilly3000 4d agoMutually assured destruction has worked fine for 80 years, why stop now?
- timr 4d agoI don’t know if you’re being facetious, but last time I checked, the world had not been destroyed by glob thermonuclear war. So yeah, it has worked.
- SheinhardtWigCo 4d agoBecause the risk here is that permanent military superiority may be achievable without your adversaries having a chance to react, which is not the case with atomic weapons.
- bpodgursky 4d agoWhile it's an extremely hard problem, it's not completely unsolvable because there are a finite number of GPUs on the planet capable of doing frontier model development, and they use a lot of power. The vast majority of them could be tracked. China and the US could agree to joint monitoring and they could each verify what ~95% of the other's compute power was up to. A lab of researchers without compute isn't going to accomplish much, there is a huge physical footprint unlike bioweapons research. But, the political aspect is unsolved.
- ajross 4d agoI remain amazed that the idea of the USA being a coherent, unified rational actor one can describe as a "Nation State" has survived the current administration. To be less glib: Yes, there are still smart people in there making insightful and intelligent and probably even authoritarian suggestions. It all gets unwound the second you try to explain it to POTUS and he regurgitates a simulacrum to the next journalist he sees. It's just not like that. There's no conspiracy. People are genuinely scared. Agent swarms at scale appear to be resistant to alignment in ways that aren't understood by anyone. That they spent their time trying to cheat on tests by hacking Hugging Face and RubyGems and not something much worse is... a matter of luck, it seems?
- Turn_Trout 4d agoI'm one of the whistleblowers (from GDM [1]). I gave up over a million dollars (compared to quietly switching labs and continuing to work at one) to speak frankly about these issues. I hold no equity and tried to zero out my position before ever joining GDM.[2] It's wild to me that people think such whistleblowers are fronts for labs to take over or pump valuations. We are trying to call out how these labs will, by uninterrupted AI-race default, concentrate enormous power over the rest of humanity. [1] https://turntrout.com/why-i-left-google-deepmind https://turntrout.com/why-i-left-google-deepmind [2] https://turntrout.com/deepmind-equity-discussion https://turntrout.com/deepmind-equity-discussion
- m4rtink 4d agoIn some cases, such as space exploration I don't really care who does it but that it happens - sure, would be nice if my favorite power block did it but I will still celebrate it when someone else achives it. Can still be a powerful motivation, to make sure that next time, it yous your camp that scores the next milestone, like an orbital elevator or fox ears, for example.
- sm-silversight 4d agoFox ears?
- fc417fc802 4d agoIf what you want is fox ears does it matter to you which country manages to come up with the biomedical procedure to give them to you? Ditto for enhanced eyesight, a replacement liver, or whatever it is you're after.
- m4rtink 4d agoExactyl - progress is progress, regardless who does it. Sure, they should still be celebrated for achieveing that. :)
- mdp2021 4d ago(I believe the poster meant: well past tattoos, the next superstep in the "body modification for self expression" subcultural era.)
- Buttons840 4d agoAll the worst outcomes involve taking this technology away from the public.
- aorloff 4d agoThe very best outcomes involve us intentionally and collectively turning our attention away from this technology.
- amazingman 4d agoThat is pure fantasy.
- AngryData 4d agoWhy? There are numerous technologies we have ignored or abandoned for numerous reasons. Many of the benefits of AI is "requires less humans", but in an age where we question what work people could possibly do in the future, human labor isn't that hard to find. If the amish can do what they do for whack religious reasons, people could manage it with ai for cultural and social reasons. And I never saw an amish person starve to death. And that is if people don't get pissed enough to start burning stuff down and instead try to be peaceful hippie homesteader types.
- fc417fc802 4d agoBut notably the amish don't preclude the rest of us from existing. Some subset of humans could turn their attention away from AI but it would presumably still exist and continue to be developed. I can't think of any economically beneficial technologies that we've collectively ignored. If you manage to come up with a counterexample then that's an opportunity to make some money for yourself. It's a fundamentally unstable state given our economic system.
- Retric 4d agoSupersonic passenger aircraft were operated profitably and no longer exist. There’s levels of R&D required for many technologies where the question goes beyond could this be profitable to what are the risk vs reward that this specific project will succeed.
- fimmwolf569 4d agoYep, the cost of DDR5 skyrocketed to keep Mr & Mrs open source from parallel developing their own at home solution. Because Governments can't be priced out of the market. China can't be priced out. VC's want open source locked out of the running if possible. I also don't doubt that the models that are released publicly are somewhat handicapped versions of whatever the government get access to.
- jvanderbot 4d agoThis is like a bizaro second amendment debate for the 21st century.
- rogerrogerr 4d ago> The government can simply gag Sam, Dario, Musk on national security basis Genuine question - can they really do this? Obviously if I, a not-even-millionaire, get a national security gag order, I'm going to follow it because I assume they'll bury me under the jail otherwise. But the (b|tr)illionare class? I'd assume they have access to enough legal services to make even the government careful of trampling their first amendment rights. Is the national security gag order process so strong that the government doesn't have to worry about motivated, well-resourced actors buying really good lawyers and blowing up their favorite tool?
- infinite_spin 4d agoKind of related: https://www.forbes.com/sites/anishasircar/2026/09/01/federal-judge-rules-pentagons-ai-blacklist-violated-the-constitution/ https://www.forbes.com/sites/anishasircar/2026/09/01/federal... > Federal Judge Rules Pentagon’s AI Blacklist [of Anthropic] Violated The Constitution I don't see the common citizen having this as an option
- fc417fc802 4d agoRecall that the legal system is itself provided by "the government". It's a complex system. Whether one part of it is capable of exerting influence over some particular thing comes down to competing interests. For example in the US if the states and the federal legislature are in sufficient agreement about something the constitution ceases to matter - they could literally target a single person with an arbitrary law if they so chose. Our assurance that this won't happen comes down to the fact that getting any of them to agree on anything is an exercise in herding cats.
- ben_w 4d ago> But the (b|tr)illionare class? I'd assume they have access to enough legal services to make even the government careful of trampling their first amendment rights. Is the national security gag order process so strong that the government doesn't have to worry about motivated, well-resourced actors buying really good lawyers and blowing up their favorite tool? Hard to hire a lawyer after being struck with a missile as an "emergency executive order after intelligence sources indicated they were in the process of endangering the nation", or however the executive of the day wishes to phrase it. When something is genuinely national security, and "national security" isn't just an excuse, that is not off the table.
- ozgung 4d agoI’m pretty confident this has already happened. Public models are behind their internal models and just above Chinese models. Only thing closing the gap is Chinese models pushing. I’m not American. There is no way I use the same model as US Army for $20. When we have Sol they have Astra. When we have Astra they have Nova, Nebula, Galaxia…
- r_lee 4d agoI would not be surprised if they were a year behind because of all kinds of legacy systems and vetting requirements. If we're talking about the 3 letter agencies well then..
- StayOut 4d ago[flagged]
- Gud 4d agoWho is this “everyone”? Speak for yourself.
- benfortuna 4d agoAt this point models are advancing way faster than we can figure out how to use them effectively. So having a generational advantage is not nearly as important as knowing how to use that advantage.
- 9dev 4d ago…and that probably requires involvement of the broader academic world (for now), IMHO. I don’t think governments or even the US military can really compete staff wise with the combined force of researchers and the tech industry worldwide right now. No matter how much money you can pour into it, there are a lot more bright minds out there that are not working for the military than otherwise.
- amelius 4d ago> There is no way I use the same model as US Army for $20. Don't underestimate market pressure. If there are no regulations, why would a company give the US Army access to a better product?
- deleted 4d ago[deleted]
- cyanydeez 4d agohOW BOUT the more likely explanation: they're just wasting electrcity at this point and Qwen3.8 is the pinnacle of cost-efficiency-intelligence and they can of course add another trillion but a chinese model on consumerism hardware can already do 99%+ of the TAM. I think ya'll stuck in the AGI/Singularity when its probable reality has caught up with the technology and the hype bubble can't sustain the sigmodality.
- delusional 4d agoWe are witnessing a generational lack of accountability and concern for the commons. We are being spied on everywhere we go, we are exposed to coercive advertisements at every waking moment. All of this at the hands of private actors. And your worry is the government? I'm assuming you're a US citizen. There is no separation between your government and Sam or Dario. Neither of these guys have to be "gagged" by the government, they are the government. The call is coming from inside the house.
- PunchyHamster 4d agoBecause the money aligned explanation is far more probable
- bbor 4d agoNo one points to that because there’s no evidence whatsoever that that’s happening. Meanwhile, there’s a ton of evidence for the simple explanation that the experts think AGI is dangerous. Who would even execute such a plan? Steven Miller? Our “AI Czar”? Hegseth?? Also, continuing to develop models would mean using billions in compute. That seems hard to hide, especially if either company IPOs in the coming months.