6 ms·
> with all of this traffic served on Chinese AI chips RIP Nivida shareholders
by sunbum 21d ago
> with all of this traffic served on Chinese AI chips
RIP Nivida shareholders
- ChoosesBarbecue 21d agoGod I wish I could’ve shorted NVIDIA right now
- browningstreet 21d agoIt's earnings day for them...
- kingstnap 21d agoWhats stopping you? You could buy puts right now. Get a 210 strike put contract and if your thesis is that nvidias current 10 day slide continues you could make some money.
- outworlder 21d agoUnless NVidia craters you are likely to lose money given the IV crush that will happen today.
- Bluestein 21d agoThis is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.- Further quote: "Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale." https://z.ai/blog/glm-5.3-flash https://z.ai/blog/glm-5.3-flash
- xtracto 21d agoAnyone knows what are those Chinese chips? Can they be bought? (Assuming im not i the US, And actually im in a 3rd world country).
- gunalx 21d agoThey are high end really expensive Huawei ascend GPUs. It is kinda bruteforcing the performance on a older semiconductor processing tech, so total production is pretty low.
- ayewo 21d agoPerhaps these may be Huawei Ascend chips. https://en.wikipedia.org/wiki/HiSilicon#Ascend_910 https://en.wikipedia.org/wiki/HiSilicon#Ascend_910 https://medium.com/@huaweiclouddevelper/a-brief-introduction-to-huawei-ascend-cloud-cbef8f25bc34 https://medium.com/@huaweiclouddevelper/a-brief-introduction...
- bean469 20d ago> comparable to mainstream NVIDIA GPUs By this they probably mean RTX series GPUs? If so, then they are not comparing the hardware efficiency with the A100 / H100, etc. that are commonly used for training models
- rvz 21d agoThis is no surprise [0] [1]. >> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators." It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish. [0] https://news.ycombinator.com/item?id=49397204 https://news.ycombinator.com/item?id=49397204 [1] https://news.ycombinator.com/item?id=49431231 https://news.ycombinator.com/item?id=49431231
- ThouYS 21d agoyay, I called it! :) (in the other thread)
- dannyw 21d agoAnother self-inflicted own courtesy of US government policy. While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.
- ignoramous 21d agoThe export controls were revoked before it triggered Chinese protectionism: https://www.silicon.co.uk/e-innovation/artificial-intelligence/china-nvidia-h200-627924 https://www.silicon.co.uk/e-innovation/artificial-intelligen... / https://archive.vn/B2pah https://archive.vn/B2pah
- mlinsey 21d agoRevoked or not, just ever having those controls signals to the Chinese ecosystem that you're not necessarily a reliable supplier (Would you trust US export policy to remain stable for the next ~decade given the state of US politic?) and to the Chinese government just how strategically important you see these components. This isn't the kind of thing you can hash out in public and go back and forth on. Once you put it out there, the other party will take steps to make sure they don't have to rely on us in the long run.
- re-thc 21d ago> The export controls were revoked before Zai is on another "export control" list outside the broader 1. Doesn't help.
- bigbadfeline 21d agoThe export controls were not revoked, only reduced, and not before, but after China refused to buy low performing chips. Top gear was and is still sanctioned, as is any EUVL equipment.
- bigbadfeline 21d agoAnd to add to the above: by building their own supply chain for chips, China is helping the unprivileged, those who can't front-run the market with long-term contracts. If China wasn't producing their own chips, the prices for us would be even higher. Similar to the war-pricing of oil, China's reduction of imports is actually helping to keep our inflation from going even higher.
- redox99 21d agoNot really a brag: it ran like shit. Very slow (~20tps, VERY high latency) and it would timeout all the time. I'm sure the chips are fine, but they clearly didn't have enough capacity for the demand they had (that 100T/day claim was asbolute bs)
- Implicated 21d ago[flagged]
- nchmy 21d agoseems unlikely that they'll get nearly as much demand now that it isnt free
- redox99 21d agoSure, although I still expect it to become the most used model on openrouter.
- Aurornis 21d agoOx Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.
- HDBaseT 21d agoOx Alpha was also serving 10T+ tokens a day for free. When it first launched on OpenRouter I was getting nearly 70 Tokens/second.
- knowaveragejoe 21d agoHas there been any confirmation about what that model even is? Edit: Ah: > This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.
- cortesoft 21d agoIt's also in this very announcement, in the first paragraph: > Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.
- VulgarExigency 21d agoIt was being served for free. They were almost certainly being overloaded.
- Aurornis 21d agoPresumably the efficiency numbers they're quoting are for the high concurrency state they were serving. RAM was probably the bottleneck for the amount of context they were offering. I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful
- Implicated 21d ago> and it was running very slowly ... I'm at a loss for words here. It was being served for free. To the entire world.
- WarmWash 21d agoI don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini) And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet. So that leaves local hosting/leasing, but one of those has totally non-practical economics and the other doesn't have enough compute to meet any kind of real demand. I also have yet to meet a single person who isn't neck-deep in the tech space mention a Chinese LLM. It's 100% the big American three. If anything it's custom chips from the labs that threatens Nvidia.
- Jcampuzano2 21d agoGenuine question but who do you put as the "three" in big three. Because I genuinely can't tell if you mean Google or SpaceX/X.ai lol.
- WarmWash 21d agoGoogle probably serves more tokens then OAI and Anthropic combined, even if many of those tokens aren't from explicit gemini requests, but from AI overviews and other service integrations. xAI is already selling spare compute, and basically exists just to gas spacex's perceived valuation.
- pianopatrick 21d agoI can easily see a situation where most non American AI usage is on Chinese models on Chinese chips though.
- uhfraid 21d agoWhat about the current situation, where serious API payers are increasingly OK with using open-weight models running on US providers? https://www.ft.com/content/32a70a3c-7d28-40b4-808e-36edb58c7d01 https://www.ft.com/content/32a70a3c-7d28-40b4-808e-36edb58c7...
- rapind 21d ago
- bityard 21d agoMost US companies that have anything to do with government, finance, medical, etc. already have contractual or regulatory obligations which prevent them from using Chinese hardware or services, even before the AI boom. That's a huge market. Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)
- cheema33 21d ago> Most US companies that have anything to do with government, finance, medical, etc... That's a huge market. Compared to the rest of the world?
- bityard 21d agoI don't have any pie charts in front of me, but yes, I would estimate it's a decently big slice of the world market.
- itemize123 21d agobigger than RotW
- saberience 21d agoNot really. Chinese AI companies were never using NVidia AI chips. This announcement doesn't really mean anything at all. It means the very few people who are already using Z.ai's API will continue to do so, but the vast majority of money going to Nvidia is through the massive amount of business going to Anthropic, OpenAI, and other western cloud providers and inference providers, who are mostly using NVidia chips for inference. Also, NVidia chips are still sold out and supply constrained.