10 ms·
Intel Gaudi 3 AI Accelerator
- 1024core 2y ago> Memory Boost for LLM Capacity Requirements: 128 gigabytes (GB) of HBMe2 memory capacity, 3.7 terabytes (TB) of memory bandwidth ... I didn't know "terabytes (TB)" was a unit of memory bandwidth...
- throwup238 2y agoIt’s equivalent to about thirteen football fields per arn if that helps.
- gnabgib 2y agoBit of an embarrassing typo, they do later qualify it as 3.7TB/s
- SteveNuts 2y agoMost of the time bandwidth is expressed in giga/gibi/tera/tebi bits per second so this is also confusing to me
- sliken 2y agoOnly for networking, not for anything measured inside a node. Disk bandwidth, cache bandwidth, and memory bandwidth is nearly always measured in bytes/sec (bandwidth), or NS/cache line or similar (which is mix of bandwidth and latency).
- deleted 2y ago[deleted]
- nahnahno 2y agoAbout as relevant a measure of speed as parsecs
- rileyphone 2y ago128GB in one chip seems important with the rise of sparse architectures like MoE. Hopefully these are competitive with Nvidia's offerings, though in the end they will be competing for the same fab space as Nvidia if I'm not mistaken.
- latchkey 2y agoAMD MI300x is 192GB.
- tucnak 2y agoWhich would be impressive had it _actually_ worked for ML workloads.
- Hugsun 2y agoDoes it not work for them? Where can I learn why?
- tucnak 2y agoJust go have a look around Github issues in their ROCm repositories on Github. A few months back the top excuse re: AMD was that we're not supposed to use their "consumer" cards, however the datacenter stuff is kosher. Well, guess what, we have purchased their datacenter card, MI50, and it's similarly screwed. Too many bugs in the kernel, kernel crashes, hangs, and the ROCm code is buggy / incomplete. When it works, it works for a short period of time, and yes HBM memory is kind of nice, but the whole thing is not worth it. Some say MI210 and MI300 are better, but it's just wishful thinking as all the bugs are in the software, kernel driver, and firmware. I have spent too many hours troubleshooting entry-level datacenter-grade Instinct cards with no recourse from AMD whatsoever to pay 10+ thousands for MI210 a couple-year old underpowered hardware, and MI300 is just unavailable. Not even from cloud providers which should be telling enough.
- latchkey 2y ago
- riskable 2y ago> Twenty-four 200 gigabit (Gb) Ethernet ports are integrated into every Intel Gaudi 3 accelerator WHAT‽ It's basically got the equivalent of a 24-port, 200-gigabit switch built into it. How does that make sense? Can you imaging stringing 24 Cat 8 cables between servers in a single rack? Wait: How do you even decide where those cables go? Do you buy 24 Gaudi 3 accelerators and run cables directly between every single one of them so they can all talk 200-gigabit ethernet to each other? Also: If you've got that many Cat 8 cables coming out the back of the thing how do you even access it? You'll have to unplug half of them (better keep track of which was connected to what port!) just to be able to grab the shell of the device in the rack. 24 ports is usually enough to take up the majority of horizontal space in the rack so maybe this thing requires a minimum of 2-4U just to use it? That would make more sense but not help in the density department. I'm imagining a lot of orders for "a gradient" of colors of cables so the data center folks wiring the things can keep track of which cable is supposed to go where.
- parentheses 2y agoRainbow parens, meet rainbow tables.
- deleted 2y ago[deleted]
- radicaldreamer 2y agoThe amount of power that will use up is massive, they should've gone for some fiber instead
- latchkey 2y ago> the only MLPerf-benchmarked alternative for LLMs on the market I hope to work on this for AMD MI300x soon. My company just got added to the MLCommons organization.
- colechristensen 2y agoAnyone have experience and suggestions for an AI accelerator? Think prototype consumer product with total cost preferably < $500, definitely less than $1000.
- hedgehog 2y agoWhat else in on the BOM? Volume? At that price you likely want to use whatever resources are on the SoC that runs the thing and work around that. Feel free to e-mail me.
- jsheard 2y agoThe default answer is to get the biggest Nvidia gaming card you can afford, prioritizing VRAM size over speed. Ideally one of the 24GB ones.
- mirekrusin 2y agoRent or 3090, maybe used 4090 if you're lucky.
- Hugsun 2y agoYou can get very cheap tesla P40s with 24gb of ram. They are much much slower than the newer cards but offer decent value for running a local chatbot. I can't speak to the ease of configuration but know that some people have used these successfully.
- jononor 2y agoWhat is the workload?
- wmf 2y agoAMD Hawk Point?
- dist-epoch 2y agoAll new CPUs will have so called NPUs inside them. For helping running models locally.
- JonChesterfield 2y ago
- neilmovva 2y agoA bit surprised that they're using HBM2e, which is what Nvidia A100 (80GB) used back in 2020. But Intel is using 8 stacks here, so Gaudi 3 achieves comparable total bandwidth (3.7TB/s) to H100 (3.4TB/s) which uses 5 stacks of HBM3. Hopefully the older HBM has better supply - HBM3 is hard to get right now! The Gaudi 3 multi-chip package also looks interesting. I see 2 central compute dies, 8 HBM die stacks, and then 6 small dies interleaved between the HBM stacks - curious to know whether those are also functional, or just structural elements for mechanical support.
- bayindirh 2y ago> A bit surprised that they're using HBM2e, which is what Nvidia A100 (80GB) used back in 2020. This is one of the secret recipes of Intel. They can use older tech and push it a little further to catch/surpass current gen tech until current gen becomes easier/cheaper to produce/acquire/integrate. They have done it with their first quad core processors by merging two dual core processors (Q6xxx series), or by creating absurdly clocked single core processors aimed at very niche market segments. We have not seen it until now, because they were sleeping at the wheel, and knocked unconscious by AMD.
- JonChesterfield 2y ago> This is one of the secret recipes of Intel Any other examples of this? I remember the secret sauce being a process advantage over the competition, exactly the opposite of making old tech outperform the state of the art.
- calaphos 2y agoIntels surprisingly fast 14nm processors come to mind. Born of necessity as they couldn't get their 10 and later 7nm processes working for years. Despite that Intel managed to keep up in single core performance with newer 7nm AMD chips, although at a mich higher power draw.
- Dalewyn 2y ago
- jacksonhacker 2y ago[dead]
- sairahul82 2y agoCan we expect the price of 'Gaudi 3 PCIe' to be reasonable enough to put in a workstation? That would be a game changer for local LLMs
- CuriouslyC 2y agoJust based on the RAM alone, let's just say if you can't just buy a Vision Pro without a second thought about the price tag, don't get your hopes up.
- wongarsu 2y agoProbably not. An 40GB Nvidia A100 is arguably reasonable for a workstation at $6000. Depending on your definition an 80GB A100 for $16000 is still reasonable. I don't see this being cheaper than an 80GB A100. Probably a good bit more expensive, seeing as it has more RAM, compares itself favorably to the H100, and has enough compelling features that it probably doesn't have to (strongly) compete on price.
- narrator 2y agoIsn't it much better to get a Mac Studio with an M2 Max and 192gb of Ram and 31 terraflops for $6599 and run llama.cpp?
- magic_hamster 2y agoMacs don't support CUDA which means all that wonderful hardware will be useless when trying to do anything with AI for at least a few years. There's Metal but it has its own set of problems, biggest one being it isn't a drop in CUDA replacement.
- doublepg23 2y agoI'm assuming this won't support CUDA either?
- adam_arthur 2y ago
- yieldcrv 2y agoHas anyone here bought an AI accelerator to run their AI SaaS service from their home to customers instead of trying to make a profit on top of OpenAI or Replicate Seems like an okay $8,000 - $30,000 investment, and bare metal server maintenance isn’t that complicated these days.
- shiftpgdn 2y agoDingboard runs off of the owner's pile of used gamer cards. The owner frequently posts about it on twitter.
- kaycebasques 2y agoWow, I very much appreciate the use of the 5 Ws and H [1] in this announcement. Thank you Intel for not subjecting my eyes to corp BS [1] https://en.wikipedia.org/wiki/Five_Ws https://en.wikipedia.org/wiki/Five_Ws
- belval 2y agoI wonder if with the advent of LLMs being able to spit out perfect corpo-speak everyone will recenter to succint and short "here's the gist" as the long version will become associated to cheap automated output.
- YetAnotherNick 2y agoSo now hardware companies stopped reporting FLOP/s number and reports in arbitrary unit of parallel operation/s.
- AnonMO 2y ago1835 tflops fp8. you have to look for it, but they posted it. The link in the op is just an announcement. the white paper has more info. https://www.intel.com/content/www/us/en/content-details/817486/intel-gaudi-3-ai-accelerator-white-paper.html https://www.intel.com/content/www/us/en/content-details/8174...
- whalesalad 2y agohttps://www.merriam-webster.com/dictionary/gaudy https://www.merriam-webster.com/dictionary/gaudy
- jagger27 2y agohttps://en.wikipedia.org/wiki/Antoni_Gaud%C3%AD https://en.wikipedia.org/wiki/Antoni_Gaud%C3%AD
- riazrizvi 2y agoThat’s an i. He’s one the the greatest architects of all time. https://www.archdaily.com/877599/10-must-see-gaudi-buildings-in-barcelona https://www.archdaily.com/877599/10-must-see-gaudi-buildings...
- TheAceOfHearts 2y agoHonestly, I thought the same thing upon reading the name. I'm aware of the reference to Antoni Gaudí, but having the name sound so close to gaudy seems a bit unfortunate. Surely they must've had better options? Then again I don't know how these sorts of names get decided anymore.
- whalesalad 2y agoto be fair intel is not known for naming things well.
- andersa 2y agoPrice?
- margaretanthony 2y ago[dead]
- mpreda 2y agoHow much does one such card cost?
- kylixz 2y agoThis is a bit snarky — but will Intel actually keep this product line alive for more than a few years? Having been bitten by building products around some of their non-x86 offerings where they killed good IP off and then failed to support it… I’m skeptical. I truly do hope it is successful so we can have some alternative accelerators.
- forkerenok 2y agoI'm not very involved in the broader topic, but isn't the shortage of hardware for AI-related workloads intense enough so as to grant them the benefit of the doubt?
- jtriangle 2y agoThe real question is, how long does it actually have to hang around really? With the way this market is going, it probably only has to be supported in earnest for a few years by which point it'll be so far obsolete that everyone who matters will have moved on.
- AnthonyMouse 2y agoWe're talking about the architecture, not the hardware model. What people want is to have a new, faster version in a few years that will run the same code written for this one. Also, hardware has a lifecycle. At some point the old hardware isn't worth running in a large scale operation because it consumes more in electricity to run 24/7 than it would cost to replace with newer hardware. But then it falls into the hands of people who aren't going to run it 24/7, like hobbyists and students, which as a manufacturer you still want to support because that's how you get people to invest their time in your stuff instead of a competitor's.
- riffic 2y agoItanic was a fun era
- cptskippy 2y agoItanium only stuck around as long as it did because they were obligated to support HP.
- AnonMO 2y agoit's crazy that Intel can't manufacture its own chips atm, but it looks like that might change in the coming years as new fabs come online.
- alecco 2y agoGaudi 3 has PCIe 4.0 (vs. H100 PCIe 5.0, so 2x the bandwidth). Probably not a deal-breaker but it's strange for Intel (of all vendors) to lag behind in PCIe.
- wmf 2y agoN5, PCIe 4.0, and HBM2e. This chip was probably delayed two years.
- alecco 2y agoGood point, it's built on TSMC while Intel is pushing to become the #2 foundry. Probably it's because Gaudi was made by an Israeli company Intel acquired in 2019 (not an internal project). Who knows. https://www.semianalysis.com/p/is-intel-back-foundry-and-product https://www.semianalysis.com/p/is-intel-back-foundry-and-pro...
- KeplerBoy 2y agoThe whitepaper says it's PCIe 5 on Gaudi 3.
- brcmthrowaway 2y agoDoes this support apple silicon?
- ancharm 2y agoIs the scheduling / bare metal software open source through OneAPI? Can a link be posted showing it if so?
- chessgecko 2y agoI feel a little misled by the speedup numbers. They are comparing lower batch size h100/200 numbers to higher batch size gaudi 3 numbers for throughput (which is heavily improved by increasing batch size). I feel like there are some inference scenarios where this is better, but its really hard to tell from the numbers in the paper.
- m3kw9 2y agoCan you run Cuda on it?
- boroboro4 2y agoNo one runs Cuda, everyone runs PyTorch. Which you can run on it.
- m3kw9 2y agoSo does it support cuda or not are are you gonna argue little things all day?
- kimixa 2y agoCUDA is a proprietary Nvidia API where the SDK license explicitly forbids use for development of apps that might run on other hardware. You do read the licenses of SDKs you use, right? Nothing but Nvidia hardware will ever "support" CUDA.
- geertj 2y agoI wonder if someone knowledgeable could comment on OneAPI vs Cuda. I feel like if Intel is going to be a serious competitor to Nvidia, both software and hardware are going to be equally important.
- ZoomerCretin 2y agoI'm not familiar with the particulars of OneAPI, but it's just a matter of rewriting CUDA kernels into OneAPI. This is pretty trivial for the vast majority of small (<5 LoC) kernels. Unlike AMD, it looks like they're serious about dogfooding their own chips, and they have a much better reputation for their driver quality.
- JonChesterfield 2y agoAll the dev work at AMD is on our own hardware. Even things like the corporate laptops are ryzen based. The first gen ryzen laptop I got was terrible but it wasn't intel. We also do things like develop ROCm on the non-qualified cards and build our tools with our tools. It would be crazy not to.
- ZoomerCretin 2y agoYes that's why I qualified "serious" dogfooding. Of course you use your hardware for your own development work, but it's clearly not enough given that showstopper driver issues are going unfixed for half a year.
- FeepingCreature 2y agoWay more than half a year. The 7900XTX came out two years ago and still hits hardware resets with Stable Diffusion.
- deleted 2y ago[deleted]
- sorenjan 2y ago
- mk_stjames 2y agoOne nice thing about this (and the new offerings from AMD) is that they will be using the "open accelerator module (OAM)" interface- which standardizes the connector that they use to put them on baseboards, similar to the SXM connections of Nvidia that use MegArray connectors to thier baseboards. With Nvidia, the SXM connection pinouts have always been held proprietary and confidential. For example, P100's and V100's have standard PCI-e lanes connected to one of the two sides of their MegArray connectors, and if you know that pinout you could literally build PCI-e cards with SXM2/3 connectors to repurpose those now obsolete chips (this has been done by one person). There are thousands, maybe tens of thousands of P100's you could pickup for literally <$50 apiece these days which technically give you more Tflops/$ than anything on the market, but they are useless because their interface was not ever made open and has not been reverse engineered openly and the OEM baseboards (Dell, Supermicro mainly) are still hideously expensive outside China. I'm one of those people who finds 'retro-super-computing' a cool hobby and thus the interfaces like OAM being open means that these devices may actually have a life for hobbyists in 8~10 years instead of being sent directly to the bins due to secret interfaces and obfuscated backplane specifications.
- JonChesterfield 2y agoI really like this side to AMD. There's a strategic call somewhere high up to bias towards collaboration with other companies. Sharing the fabric specifications with broadcom was an amazing thing to see. It's not out of the question that we'll see single chips with chiplets made by different companies attached together.
- 01HNNWZ0MV43FF 2y agoMaybe they feel threatened by ARM on mobile and Intel on desktop / server. Companies that think they're first try to monopolize. Companies that think they're second try to cooperate.
- rhelz 2y agoWell, lets not forget, AMD is AMD because they reverse-engineered Intel chips....
- throwaway4good 2y agoWorth noting that it is fabbed by TSMC.
- amelius 2y agoMissing in these pictures are the thermal management solutions.
- InitEnabler 2y agoIf you look at one of the pictures you can get a peak at what they look like (I think...) in the bottom right. https://www.intel.com/content/dam/www/central-libraries/us/en/images/2024-04/newsroom-intel-gaudi-3-5.jpg.rendition.intel.web.1280.720.jpg https://www.intel.com/content/dam/www/central-libraries/us/e...
- wmf 2y agoIt's going to look very similar to an Nvidia SXM or AMD MI300 heatsink since these all have similar form factors.
- KeplerBoy 2y agovector floating point performance comes in at 14 Tflops/s for FP32 and 28 Tflop/s for FP16. Not the best of times for stuff that doesn't fit matrix processing units.
- einpoklum 2y agoIf your metric is memory bandwidth or memory size, then this announcement gives you some concrete information. But - suppose my metric for performance is matrix-multiply-add (or just matrix-multiply) bandwidth. What MMA primitives does Gaudi offer (i.e. type combinations and matrix dimension combinations), and how many of such ops per second, in practice? The linked page says "64,000 in parallel", but that does not actually tell me much.
- InvestorType 2y agoThis appears to be manufactured by TSMC (or Samsung). The press release says it will use a 5nm process, which is not on Intel's roadmap. "The Intel Gaudi 3 accelerator, architected for efficient large-scale AI compute, is manufactured on a 5 nanometer (nm) process"
- ac29 2y agoHabana was an acquisition and their use of TSMC predates the acquisition.
- modeless 2y agoYeah, but if Intel can't even get internal customers to adopt their foundry services it seems to bode poorly for the future of the company.
- ksec 2y agoThe design and decision to make it Fab with TSMC was way ahead of Intel's Foundry services offering. ( And it is not like Intel had the extra capacity planned at the time for Intel's GPU )
- simpsond 2y agoProcess matters. Intel was ahead for a long time, and has been behind for a long time. Perhaps they will be ahead again, but maybe not. I’d rather see them competitive.
- metadat 2y ago> Twenty-four 200 gigabit (Gb) Ethernet ports are integrated into every Intel Gaudi 3 accelerator How much does a single 200Gbit active (or inactive) fiber cable cost? Probably thousands of dollars.. making even the cabling for each card Very Expensive. Nevermind the network switches themselves.. Simultaneously impressive and disappointing.
- lillecarl 2y agohttps://www.fs.com/de-en/products/115636.html https://www.fs.com/de-en/products/115636.html 2 meters seems to be about 100$, which isn't unreasonable. If you're going fiber instead of twinax it's another order of magnitude and a bit for trancievers, but cables are pretty cheap still. You seem to be loading negative energy into this release from the get-go
- metadat 2y agoYou're going to need a lot more than 2 meters... It's probably AOC (Active-Optical Fiber Cable), they're pricey even for 40Gbit, at DC lengths.
- pezezin 2y ago2 meters is enough to connect a server to a leaf ToR switch. Now, connecting the leaf switches to the spine is a different story...
- throwaway2037 2y agoWhat do you mean by active vs inactive fiber cable? I tried to Google about this distinction, but I couldn't find anything helpful.
- metadat 2y agoMy off-the-cuff take: AOC's are a specific kind of fiber optic cable, typically used in data center applications for 100Gbit+ connections. The alternate types of fiber are typically referred to as passive fiber cables, e.g. simplex or duplex, single-mode (single fiber strands, usually in a yellow jacket) or multi-mode (multiple fiber strands, usually in an orange jacket). Each type of passive fiber cable has specific applications and requires matching transceivers, whereas AOCs are self-contained with the transceivers pre-terminated on. If you search for "AOC Fiber", lots of resources will pop up. FS.com is one helpful resource. https://community.fs.com/article/active-optical-cable-aoc-rising-star-of-telecommunications-datacom-transceiver-markets.html https://community.fs.com/article/active-optical-cable-aoc-ri... > Active optical cable (AOC) can be defined as an optical fiber jumper cable terminated with optical transceivers on both ends. It uses electrical-to-optical conversion on the cable ends to improve speed and distance performance of the cable without sacrificing compatibility with standard electrical interfaces.
- cavisne 2y agoIs there an equivalent to this reference for Intel Gaudi? https://docs.nvidia.com/cuda/parallel-thread-execution/index.html# https://docs.nvidia.com/cuda/parallel-thread-execution/index...
- sandGorgon 2y ago>Intel Gaudi software integrates the PyTorch framework and provides optimized Hugging Face community-based models – the most-common AI framework for GenAI developers today. This allows GenAI developers to operate at a high abstraction level for ease of use and productivity and ease of model porting across hardware types. what is the programming interface here ? this is not CUDA right ...so how is this being done ?
- wmf 2y agoPyTorch has a bunch of backends including CUDA, ROCm, OneAPI, etc.
- sandGorgon 2y agoi understand. but which backend is intel committing to ? not CUDA for sure. or have they created a new backend
- singhrac 2y agoIntel makes oneAPI. They have corresponding toolkits to cuDNN like oneMKL, oneDNN, etc. However the Gaudi chips are built on top of SynapseAI, another API from before the Habana acquisition. I don’t know if there’s a plan to support oneAPI on Gaudi, but it doesn’t look like it at the moment.
- MrYellowP 2y agohttps://www.dwds.de/wb/Gaudi https://www.dwds.de/wb/Gaudi That's amusing. :D