6 ms·
There’s a lot of indications that we’re currently brute forcing these models. There’s honestly not a reason they have to be 1T parameters and cost an insane amo
by lettergram 10mo ago
There’s a lot of indications that we’re currently brute forcing these models. There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference.
What we’re going to see is as energy becomes a problem; they’ll simply shift to more effective and efficient architectures on both physical hardware and model design. I suspect they can also simply charge more for the service, which reduces usage for senseless applications.
- yanhangyhy 10mo agoThere are also elements of stock price hype and geopolitical competition involved. The major U.S. tech giants are all tied to the same bandwagon — they have to maintain this cycle: buy chips → build data centers → release new models → buy more chips. It might only stop once the electricity problem becomes truly unsustainable. Of course, I don’t fully understand the specific situation in the U.S., but I even feel that one day they might flee the U.S. altogether and move to the Middle East to secure resources.
- MallocVoidstar 10mo ago> What we’re going to see is as energy becomes a problem This is much more likely to be an issue in the US than in China. https://fortune.com/2025/08/14/data-centers-china-grid-us-infrastructure/ https://fortune.com/2025/08/14/data-centers-china-grid-us-in...
- thesmtsolver 10mo agoDisagree. Part of the reason China produces more power (and pollution) is due to China manufacturing for the US. https://www.brookings.edu/articles/how-do-china-and-america-think-about-the-energy-transition/ https://www.brookings.edu/articles/how-do-china-and-america-... The source for China's energy is more fragile than that of the US. > Coal is by far China’s largest energy source, while the United States has a more balanced energy system, running on roughly one-third oil, one-third natural gas, and one-third other sources, including coal, nuclear, hydroelectricity, and other renewables. Also, China's GDP is a bit less inefficient in terms of power used per unit of GDP. China relies on coal and imports. > However, China uses roughly 20% more energy per unit of GDP than the United States. Remember, China still suffers from blackouts due to manufacturing demand not matching supply. The fortune article seems like a fluff piece. https://www.npr.org/2021/10/01/1042209223/why-covid-is-affecting-chinas-power-rations https://www.npr.org/2021/10/01/1042209223/why-covid-is-affec... https://www.bbc.com/news/business-58733193 https://www.bbc.com/news/business-58733193
- mullingitover 10mo agoThese stories are from 2021. China has been adding something like a 1GW coal plant’s worth of solar generation every eight hours in the past year, and the rate is accelerating. The US is no longer a serious competitor for China when it comes to energy production.
- deleted 10mo ago[deleted]
- DeH40 10mo agoThe reason it happened in 2021, I think, might be that China took on the production capacity gap caused by COVID shutdowns in other parts of the world. The short-term surge in production led to a temporary imbalance in the supply and demand of electricity
- timlarshanson 10mo agoThis was very surprising to me, so I just fact-check this statement (using Kimi K2 thinking, natch), and it's presently is off by a factor of 2 - 4. In 2024 China installed 277 GW solar, so 0.25 GW / 8 hours. First half of 2025 they installed 210 GW, so 0.39 GW / 8 hours. Not quite at 1 GW / 8 hrs, but approaching that figure rapidly! (I'm not sure where the coal plant comes in - really, those numbers should be derated relative to a coal plant, which can run 24/7)
- simonw 10mo ago> There’s honestly not a reason they have to be 1T parameters and cost an insane amount to train and run on inference. Kimi K2 Thinking is rumored to have cost $4.6m to train - according to "a source familiar with the matter": https://www.cnbc.com/2025/11/06/alibaba-backed-moonshot-releases-new-ai-model-kimi-k2-thinking.html https://www.cnbc.com/2025/11/06/alibaba-backed-moonshot-rele... I think the most interesting recent Chinese model may be MiniMax M2, which is just 200B parameters but benchmarks close to Sonnet 4, at least for coding. That's small enough to run well on ~$5,000 of hardware, as opposed to the 1T models which require vastly more expensive machines.
- electroglyph 10mo agoi assume that $4.6 mil is just the cost of the electricity?
- simonw 10mo agoHard to be sure because the source of that information isn't known, but generally when people talk about training costs like this they include more than just the electricity but exclude staffing costs. Other reported training costs tend to include rental of the cloud hardware (or equivalent if the hardware is owned by the company), e.g. NVIDIA H100s are sometimes priced out in cost-per-hour.
- Der_Einzige 10mo agoCitation needed on "generally when people talk about training costs like this they include more than just the electricity but exclude staffing costs". It would be simply wrong to exclude the staffing costs. When each engineer costs well over 1 million USD in total costs year over year, you sure as hell account for them.
- vanviegen 10mo agoNo, because what people are generally trying to express with numbers like these, is how much compute went into training. Perhaps another measure, like zettaflop or something would have made more sense.
- Leynos 10mo agoHaving larger models is nice because they have a much wider sphere of knowledge to draw on. Not in the sense of using them as encyclopedias. More in the sense that I want a model that is going to be able to cross reference from multiple domains that I might not have considered when trying to solve a problem.