6 ms·
Nvidia pursues $30B custom chip opportunity with new unit
- bluerooibos 3y agoI read about a new approach for making AI chips sometime last year - analogue chips by this company - https://mythic.ai/ https://mythic.ai/ Haven't heard anything about it since though.
- varelse 3y ago[dead]
- tmaly 3y agoI would love to see a consumer graphics card with 128GB VRAM Would be nice to be able to work with some of the larger open source LLM models.
- datameta 3y agoI think that goal is fine and good, but I would rather see huge investments toward in-memory compute like ReRAM and such. If we bridge the efficiency advancements of TinyML with the leap of LLM abilities, perhaps we can start on the road of not being limited by the impact of training on climate.
- asicsarecool 3y agoWhy the fuck was this downvoted. Very occasionally I get the feeling HN is entering the /. phase
- wmf 3y agoI didn't downvote it but in-memory compute is crackpot and alternative memory tech is really crackpot. It's not going to happen and it's ridiculous to propose it on the same level as GPUs with more RAM.
- zozbot234 3y agoGPU's are already closer to "in-memory compute" compared to CPU's. It's just taking the existing pattern of NUMA (non-uniform memory access) to a greater extent, as a principled approach to the so-called 'Von Neumann bottleneck'.
- imtringued 3y agoIn-memory compute is very easy, you just don't have to fall for the pipe dream of using the same process for both the memory and the compute. All you have to do is follow a package on package strategy like we already do with smartphones. A Raspberry PI 5 gets 25GB/s memory bandwidth and it only has a single DRAM chip if I recall correctly. So if you had a DIMM with 16 of these chips, you would already be on the same bandwidth as HBM. 96 DIMMs and you get 40TB/s memory bandwidth.
- deleted 3y ago[deleted]
- datameta 3y agoCan you explain further why you believe it isn't worth exploring? Is it that you can't imagine how it would scale to the level of today's high performance compute hardware? Just ten years back, squeezing ML models onto microcontrollers sounded completely insane, given their tight memory and power constraints. We've seen NN compilers developed, game-changing techniques like quantization, pruning, and graph-level optimizations pruning. This allowed deployment of ML models in microcontrollers with a newly developed framework like TFLite Micro.
- wmf 3y agoThese technologies have already been explored and they failed every time. It's throwing good money after bad. Also, speculative basic research isn't comparable to adding more RAM to a graphics card. You can't substitute one for the other.
- epistasis 3y agoAgree with the comment, but riffing on the last few words: Climate is pretty much my #1 concern about the world, but LLM use of energy is really really far down on the list of important actions for climate. First and foremost are removing roadblocks for deploying existing technologies for clean energy, and speeding up the necessary supporting infrastructure such as transmission and market policies for choosing cheapest possible solutions (over the objections of dinosaur execs that choose last century's solutions). Then the big hard to decarbonize parts of industry like cement and steel, as well as deploying electrolyzers to get ammonia fertilizer production switched over to carbon neutral production rather than from fossil-generated hydrogen. Reducing energy consumption is important for advancing AI in general, but ultimately all its energy consumption will be from clean energy sources anyway, and the switch that needs to happen is that switch in energy sources. Reducing energy use by 2x or 10x is not good enough, we must change the sources fundamentally.
- amelius 3y agoThat's overkill for any type of graphics application though. What you want is a DL or parallel compute card, not a graphics card. They are far more expensive though because compute doesn't sell to the average consumer like graphics does.
- sigmoid10 3y agoMeh, that's like saying 640kb of RAM is more than anyone would ever need. Demand follows hardware development, which in turn accelerates demand. I'm sure game developers would easily find a way to use 128GB of VRAM if it was commonly available in their target market.
- late2part 3y agoTo be fair, 640kb is more than anyone really needs. It's just far, far less than we want.
- Takennickname 3y agoCan you imagine the download sizes?
- amelius 3y agoTurns out that demand for graphics memory got stuck at a point where compute is still hungry for more. That may certainly change but it doesn't help compute much, today.
- LoganDark 3y agoIt's supply that's down - Nvidia is strictly enforcing their artificial segmentation where the jump from 24GB to 40GB of VRAM multiplies the price by around 5 (!). They know that if you need a 40GB card, you probably can't use a 24GB card at all because it'd run out of VRAM. They charge such insane prices for those higher capacity cards, because they know that they can make you pay thousands upon thousands for every tiny scrap of extra VRAM, and you will have literally no choice but to pay that price because you need the VRAM. Demand isn't low, demand is actually so high that they don't care if lowly consumers want VRAM, they already have enterprise customers that need it so badly that they'll pay prices higher than most consumers would ever imagine. And that's why they're so stingy with VRAM on the consumer cards. They need to be careful not to accidentally make them useful in the datacenter, so they can ensure that their enterprise customers continue to be forced to pay those extortionary prices. You see, it's all about the money, and it always will be. Capitalism, baby!
- Aurornis 3y agoUnfortunately, as soon as you make a card with specifications that make it great at enterprise-grade tasks, it will be bought in mass quantities by people building out data centers. This pushes the price up, as we’ve already seen. So labeling it “consumer” doesn’t really mean much. They’ve tried to enforce the distinction with EULAs before, but that doesn’t work well.
- sydd 3y ago> it will be bought in mass quantities by people building out data centers. Isn't this good? The production at scale effects will kick in, lowering the price and supply will meet demand after some hiccups.
- Aurornis 3y agoIf demand drove prices down like that then we’d already have cheap cards available. Demand puts upward pressure on prices. Supply is already maxed out and growing as fast as possible.
- piva00 3y ago> Isn't this good? The production at scale effects will kick in, lowering the price and supply will meet demand after some hiccups. You didn't account for the "hiccups", which can vary from 5-20 years until competition catches up, longer than the life of many companies. In spherical cows worlds of economics that would be just a hiccup.
- LoganDark 3y ago> Supply is already maxed out and growing as fast as possible. No, it's not anywhere near maxed out anymore. The chip shortage has been over for a while; now there is too much supply. But that's not the issue. The issue is that Nvidia really wants to force these enterprise customers to pay extremely high prices. This is why they are so afraid to give any more VRAM to the consumer chips, because if they were at all suitable for VRAM-heavy workloads, every HPC company would buy a couple $2,000 consumer cards in place of each $15,000 datacenter card. That'd lose Nvidia something like 70% which they would find entirely unacceptable.
- karolist 3y agoNot exactly what you've asked but Mac Studio exists, with 192GB at that.
- rrrix1 3y agoCalling a Mac Studio equipped with 192GB of RAM "consumer grade" is a big stretch. "Consumer ${thing}" to me is somewhere =< $3000.
- karolist 3y agoTrue, but you get a full general compute machine for under $6k, after a few years you can sell it for 3-4k and upgrade to the latest one. Personally I think it's the best thing to happen to local LLMs, by accident. I have an M1 Max 64GB and I'm blown away every day by these local models. Apple didn't plan for it, but it happened by accident due to unified memory with godzillion throughput having integrated silicon.
- brucethemoose2 3y agoWe are getting that (in early 2025?) With AMD Strix Halo. 40 CUs, 256 bit LPDDR5X, 16 CPU cores. Or so the rumors say.
- blackoil 3y agoAs a side effect we are seeing lot of investment/innovation in 2-13B models. Considering price of 4090, 128GB GPU will be >6000 which practically no one can afford.
- elabajaba 3y agoThis isn't possible currently unless you use HBM, which is significantly more expensive both to actually buy the HBM, and to package it (and all the capacity for packaging HBM interposers is already used for data center hardware). The most VRAM you could have on a consumer GPU today is 48GB, and that'd be on a 4090 or 7900xtx with clamshell VRAM (which increases cost and makes cooling significantly harder due to putting GDDR6 chips on the back of the GPU where there aren't any fans). To calculate how much VRAM is possible, you just need to divide the GPU's bus width by 32 (or 64 for Samsung's new weird double capacity but double bit width) and multiply that by the largest GDDR capacity currently available (16Gbit). As for why GPUs don't increase their bus width, there have been GPUs with 512bit busses in the past, but it makes it quite a bit more expensive (more vram chips, more traces to run, might require more/heavier PCB layers) and increases power draw.
- jedberg 3y agoJust today I was reading the article about OpenAI wanting $7T to develop their own AI chips. In the comments were a bunch of people talking about all the startups in the last 18 months trying to make bespoke AI chips. This makes a lot of sense for NVIDIA. They have the expertise, the money, the scale, and the experience already. They can probably do it cheaper than any startup and then either pass on that savings or make more profit.
- jetbalsa 3y agoDon't forget the tooling, ROCm still hasn't taken off very well.
- zozbot234 3y agoThe Mesa folks are working on the tooling situation with RustiCL, which has potential to support SYCL too in the future, and quite possibly HIP. Not just for ROCm-supported devices too, but across the board (subject to pure hardware constraints).
- roenxi 3y agoROCm runs PyTorch and TensorFlow. It seems to have more or less caught up on the technical capability front. There are outstanding problems, particularly I've found it very crash prone on a consumer desktop and wouldn't recommend an AMD card for research compute tasks where you are also running an X server using the same card. But there aren't $30 billion opportunities for custom chips on the consumer desktop right now - I'm guessing these will be for SaaS businesses where AMD are focusing. IE, it won't matter that they can't X.org and multiply matrices at the same time because servers won't use the cards for graphics.
- imtringued 3y agoPeople don't seem to understand that running neural network inference is very easy. It's not the machine learning frameworks and libraries that are difficult to get right. Those are the trivial part. The hard part is getting a culture that gives a damn about developing software that works and designing the hardware to support the features that the software needs. AMD has not figured out how to run both graphics and compute on the same GPU. There can be many reasons for that, but honestly it is probably because they either don't have the necessary virtualization hardware or because two different drivers are conflicting with one another.
- m3kw9 3y agoImagine if they got ARM, sort of good they did not as the competition would suffer
- jprd 3y agoNvidia isn't in the Fab biz, so maybe this will be easier for them to generate customer interest in a way that Intel has not been able to?
- wmf 3y agoI wonder if customers really want custom chips or just cheaper ones. Many of these custom AI chips are slower than flagship GPUs so presumably a cut-down GPU at a lower price would be just as good.
- brucethemoose2 3y agoThere is a lot of silicon consumer GPUs don't "need" for AI. But on the other hand the software stack is very mature and they are heavily amortized by the huge volume, so its kinda hard to argue with. In fact its so good that Nvidia can charge outrageous prices for the L40, A10 and such and then turn around and sell the exact same dies to consumers (with less memory).
- tomasGiden 3y agoFor customers like Ericsson it wouldn’t surprise me if they request special instructions and special hard function blocks. In telecom there are certain operations that’s specified by the standard (and some which aren’t but used as a de facto standard) which are performed so often that you want to do them in hardware instead of in software. Or the opposite, Ericsson just wants to integrate NVIDIAs IP into Ericsson’s own ASICs instead of using their own cores and other third party cores.
- wslh 3y agoI imagine there will be cheaper service providers soon for training (2024/2025). Like what companies such as Hetzner, Digital Ocean and others are providing for cloud. They are not in the same league of AWS, Google Cloud, Azure but can add more specific cloud services.
- brucethemoose2 3y agoAWS/Azure prices are really awful TBH. There are already much better places to get GPUs.
- wslh 3y agoI know but, for example, Google Cloud has a current advantage with their own hardware (TPUs). What is approximately the cost of training something like ChatGPT or Gemini? They have an advantage because they can rely in Azure and Google respectively without paying anything or with subsidised prices. Could a new player compete with them for training for other companies?
- brucethemoose2 3y agoGoogle prices the TPUs pretty exorbitantly, actually. But they give a lot of TPU time away for research, which is nice. It seems Intel Gaudi 2 is priced in a sweeter spot, but I've never head of anyone but Intel using them.
- wslh 3y agoThen I understand their advantage is training their own models and pricing it high for others no matter the cost.
- zerreh50 3y agoFrom Nvidia's history of working with AIBs, Sony, Apple, the Linux community, and probably many more, they seem to be a very hard company to work with. They have an idea of what the product looks like and it's their way or the highway. I wonder if this new department will change that. If it doesn't, it won't amount to much.
- kevingadd 3y agoTo be fair, it seems like all the game consoles that shipped with NV chips in them have been fairly successful, or their failures were explicitly not related to the parts of the silicon that NV designed. Other than perhaps the jailbreak issues the launch Switch had... And Geforce mobile parts are still quite popular in laptops. It makes me wonder how much of the difficulty is "hard company to work with" and how much is just "weird constraints make integration a pain" - I still don't know if they don't want to open their drivers, or if they can't for IP reasons.
- ksec 3y ago> all the game consoles For Nintendo Switch. they picked it up was mostly due to Nvidia being very desperate for a win, willing to sell it for cheap, using older technology and node all while having very little driver support. ( Also I remember Jensen loves Nintendo ) > I still don't know if they don't want to open their drivers Drivers for GPU is pretty much like CUDA for GPGPU. It is where 99% of the value comes from.
- aurareturn 3y agoI don't think AIBs had a hard time working with Nvidia. If you're referring to Evga, they just wanted to exit the GPU business after the crypto boom. For Sony, I don't recall any problems. I think Sony and Microsoft just wanted a supplier that can provide an APU and only AMD could do it at the time. AMD was also on the verge of bankruptcy so they gave Sony and Microsoft favorable terms. For Apple, it was a case who was to blame for the GPU failures. Ultimately, I think it would have been better for Macs to use Nvidia GPUs over AMD GPUs. Not sure about Linux. But I assume most big Nvidia server deployments run on Linux. I personally think Nvidia did not do much custom chip is because the margins weren't there and they wanted to devote their resources to AI. Obviously they were correct in their choice. Their market cap is closing in on Apple. Let that sink it for a bit. The only exception is the Nintendo Switch and I have a feeling Nvidia just wanted to be in at least one gaming device to say they're still in it.
- carld 3y agoIs TSMC the exclusive manufacturer for this unit? I can’t find the info.
- paulmd 3y agoit wouldn't be uncommon for the fab not to be announced or not announced up front. it's not really a consumer-facing spec, and it's not NVIDIA's announcement to make. samsung and intel foundry services are both serious possibilities imo, NVIDIA is incredibly portable across nodes and will take advantage of anything that's cheap and makes sense as a product. for example pascal used samsung 14LPP for the 1050 and 1050 Ti, and the A100 was taped out on TSMC 7nm (not samsung 8nm), etc. They are actually quite diversified, they have a product foothold on almost every node that matters, if Samsung suddenly becomes a blazing deal they're ready to go. Etc. they also already signed a semicustom licensing deal with Mediatek last year, this is part of an overall trend of NVIDIA pivoting towards licensing and platform. ARM wasn't a play to sell more Tegras, it was a play to have GeForce be the default IP for the base tier ARM licenses. https://corp.mediatek.com/news-events/press-releases/mediatek-partners-with-nvidia-to-provide-full-scale-product-roadmap-to-the-automotive-industry https://corp.mediatek.com/news-events/press-releases/mediate... in a world where software innovation is replacing hardware innovation... platform is king. He's struck gold, now he is trying to convert it into platform. And with Sony pivoting towards AI/ML and RT upscaling with PS5 Pro, they will be essentially on par with Ada in broader feature set. And the Playstation API is a platform worthy to rival his own, Sony is uniquely positioned to go after him in the AI market and leverage studios into creating some value for them (especially with CPU speeds not increasing - use the tensors or don't, I guess!). Apple is surging ahead with Metal too - they have an excellent platform as well. His time is not unlimited here. https://www.resetera.com/threads/tom-henderson-ps5-pro-specs-and-release-window-details-codenamed-trinity-30wgps-18000mts-memory-speed-november-2024-target.744703/page-63?post=116078280#post-116078280 https://www.resetera.com/threads/tom-henderson-ps5-pro-specs...
- ksec 3y agoI would assume instead of Google and Microsoft designing their own ARM Processor on the server using ARM's IP with custom design. Nvidia could do that for them. After all Microsoft and Google dont have the economy of scale, nor do they have the ( or as much ) expertise. Nvidia could also provide other IPs such as Network, Ethernet and other GPGPU integration. I guess that is why they are in talks with Ericsson. Ericsson has a history of working with Intel and Intel failed to deliver. Basically I think Nvidia will now move all the sunk cost of New Node to every new generation of AI chip. While the rest of the business unit, from GPU, Network, SoC, and now Custom Chips benefits from it. I am even wondering if Nvidia will come back to Smartphone Mobile SoC.
- osigurdson 3y agoI wonder if we are somehow at peak GPU profitability? It seems that either efficiencies in AI or competition will emerge.
- andy81 3y agoDiminishing returns on more pixels is also a factor. 4k vs 1440 is 2.4x the pixels (and related compute/heat/energy), but a barely-visible difference in visual clarity. We can see successful consoles like the Switch that have given up on competing on resolution already.
- alecco 3y agohttps://archive.ph/TtIcY https://archive.ph/TtIcY