10 ms·
I feel this is bigger than the 5x series GPUs. Given the craze around AI/LLMs, this can also potentially eat into Apple’s slice of the enthusiast AI dev segment
by Karupan 2y ago
I feel this is bigger than the 5x series GPUs. Given the craze around AI/LLMs, this can also potentially eat into Apple’s slice of the enthusiast AI dev segment once the M4 Max/Ultra Mac minis are released. I sure wished I held some Nvidia stocks, they seem to be doing everything right in the last few years!
- iKevinShah 2y agoI can confirm this is the case (for me).
- dagmx 2y agoI think the enthusiast side of things is a negligible part of the market. That said, enthusiasts do help drive a lot of the improvements to the tech stack so if they start using this, it’ll entrench NVIDIA even more.
- option 2y agotoday’s enthusiast, grad student, hacker is tomorrow’s startup founder, CEO, CTO or 10x contributor in large tech company
- Mistletoe 2y ago> tomorrow’s startup founder, CEO, CTO or 10x contributor in large tech company Do we need more of those? We need plumbers and people that know how to build houses. We are completely full on founders and executives.
- davrosthedalek 2y agoWe might not, but Nvidia would certainly like it.
- hatboat 2y agoIf they're already an "enthusiast, grad student, hacker", are they likely to choose the "plumbers and people that know how to build houses" career track? True passion for one's career is rare, despite the clichéd platitudes ecouraging otherwise. That's something we should encourage and invest in regardless of the field.
- computably 2y agoYeah, it's more about preempting competitors from attracting any ecosystem development than the revenue itself.
- VikingCoder 2y agoIf I were NVidia, I would be throwing everything I could at making entertainment experiences that need one of these to run... I mean, this is awfully close to being "Her" in a box, right?
- dagmx 2y agoI feel like a lot of people miss that Her was a dystopian future, not an ideal to hit. Also, it’s $3000. For that you could buy subscriptions to OpenAI etc and have the dystopian partner everywhere you go.
- tacticus 2y agothey don't miss that part. they just want to be the evil character.
- OhioMan2943 2y agoThe dystopian overton window has shifted, didn't you know, moral ambiguity is a win now? :) Tesla was right.
- VikingCoder 2y agoWe already live in dystopian hell and I'd like to have Scarlett Johansen whispering in my ear, thanks. Also, I don't particularly want my data to be processed by anyone else.
- croes 2y agoOpenAI doesn’t make any profit. So either it dies or prices go up. Not to mention the privacy aspect of your own machine and the freedom of choice which models to run
- blackoil 2y ago> So either it dies or prices go up. Or efficiency gains in hardware and software catchup making current price point profitable.
- qwertox 2y agoYou could have said the same about gamers buying expensive hardware in the 00's. It's what made Nvidia big.
- Cumpiler69 2y agoThere's a lot more gamers than people wanting to play with LLms at home.
- anonylizard 2y agoThere's a titanic market with people wanting some uncensored local LLM/image/video generation model. This market extremely overlaps with gamers today, but will grow exponentially every year.
- Cumpiler69 2y agoHow big is that market you claim? Local LLM image generation already exists out off the box on latest Samsung flagship phones and it's mostly a Gimmick that gets old pretty quickly. Hardly comparable to gaming in terms of market size and profitablity. Plus, YouTube and the Google images is already full of AI generated slop and people are already tired of it. "AI fatigue" amongst majority of general consumers is a documented thing. Gaming fatigues is not.
- madwolf 2y agoI think he implied AI generated porn. Perhaps also other kind of images that are at odds with morality and/or the law. I'm not sure but probably Samsung phones don't let you do that.
- TeMPOraL 2y ago> Gaming fatigues is not. It is. You may know it as the "I prefer to play board games (and feel smugly superior about it) because they're ${more social, require imagination, $whatever}" crowd.
- Karupan 2y agoI’m not so sure it’s negligible. My anecdotal experience is that since Apple Silicon chips were found to be “ok” enough to run inference with MLX, more non-technical people in my circle have asked me how they can run LLMs on their macs. Surely a smaller market than gamers or datacenters for sure.
- dagmx 2y agoI mean negligible to their bottom line. There may be tons of units bought or not, but the margin on a single datacenter system would buy tens of these. It’s purely an ecosystem play imho. It benefits the kind of people who will go on to make potentially cool things and will stay loyal.
- htrp 2y ago>It’s purely an ecosystem play imho. It benefits the kind of people who will go on to make potentially cool things and will stay loyal. 100% The people who prototype on a 3k workstation will also be the people who decide how to architect for a 3k GPU buildout for model training.
- mrlongroots 2y ago> It’s purely an ecosystem play imho. It benefits the kind of people who will go on to make potentially cool things and will stay loyal. It will be massive for research labs. Most academics have to jump through a lot of hoops to get to play with not just CUDA, but also GPUDirect/RDMA/Infiniband etc. If you get older/donated hardware, you may have a large cluster but not newer features.
- ckemere 2y agoAcademic minimal-bureaucracy purchasing card limit is about $4k, so pricing is convenient*2.
- bwfan123 2y agoDevalapers developers developers - balmer monkey dance - the key to be entrenched is the platform ecosystem. Also why aws is giving trainium credits for free
- gr3ml1n 2y agoAMD thought the enthusiast side of things was a negligible side of the market.
- dagmx 2y agoThat’s not what I’m saying. I’m saying that the people buying this aren’t going to shift their bottom line in any kind of noticeable way. They’re already sold out of their money makers. This is just an entrenchment opportunity.
- epolanski 2y agoIf this is gonna be widely used by ML engineers, in biopharma, etc and they land 1000$ margins at half a million sales that's half a billion in revenue, with potential to grow.
- paxys 2y ago“Bigger” in what sense? For AI? Sure, because this an AI product. 5x series are gaming cards.
- Karupan 2y agoBigger in the sense of the announcements.
- Auracle 2y agoEh. Gaming cards, but also significantly faster. If the model fits in the VRAM the 5090 is a much better buy.
- a________d 2y agoNot expecting this to compete with the 5x series in terms of gaming; But it's interesting to note the increase in gaming performance Jensen was speaking about with Blackwell was larger related to inferenced frames generated by the tensor cores. I wonder how it would go as a productivity/tinkering/gaming rig? Could a GPU potentially be stacked in the same way an additional Digit can?
- wpwpwpw 2y agoWould hadn't nvidia cripple nvlink on geforce.
- qwertox 2y agoThis is somewhat similar to what GeForce was to gamers back in the days, but for AI enthusiasts. Sure, the price is much higher, but at least it's a completely integrated solution.
- Karupan 2y agoYep that's what I'm thinking as well. I was going to buy a 5090 mainly to play around with LLM code generation, but this is a worthy option for roughly the same price as building a new PC with a 5090.
- qwertox 2y agoIt has 128 GB of unified RAM. It will not be as fast as the 32 GB VRAM of the 5090, but what gamer cards have always lacked was memory. Plus you have fast interconnects, if you want to stack them. I was somewhat attracted by the Jetson AGX Orin with 64 GB RAM, but this one is a no-brainer for me, as long as idle power is reasonable.
- deleted 2y ago[deleted]
- moffkalast 2y agoHaving your main pc as an LLM rig also really sucks for multitasking, since if you want to keep a model loaded to use it when needed, it means you have zero resources left to do anything else. GPU memory maxed out, most of the RAM used. Having a dedicated machine even if it's slower is a lot more practical imo, since you can actually do other things while it generates instead of having to sit there and wait, not being able to do anything else.
- puppymaster 2y agoit eats into all NVDA consumer-facing clients no? I can see why openai and etc are looking for alternative hardware solution to train their next model.
- doctorpangloss 2y agoWhat slice? Also, macOS devices are not very good inference solutions. They are just believed to be by diehards. I don't think Digits will perform well either. If NVIDIA wanted you to have good performance on a budget, it would ship NVLink on the 5090.
- YetAnotherNick 2y ago> Also, macOS devices are not very good inference solutions They are good for single batch inference and have very good tok/sec/user. ollama works perfectly in mac.
- Karupan 2y agoThey are perfectly fine for certain people. I can run Qwen-2.5-coder 14B on my M2 Max MacBook Pro with 32gb at ~16 tok/sec. At least in my circle, people are budget conscious and would prefer using existing devices rather than pay for subscriptions where possible. And we know why they won't ship NVLink anymore on prosumer GPUs: they control almost the entire segment and why give more away for free? Good for the company and investors, bad for us consumers.
- acchow 2y ago> I can run Qwen-2.5-coder 14B on my M2 Max MacBook Pro with 32gb at ~16 tok/sec. At least in my circle, people are budget conscious Qwen 2.5 32B on openrouter is $0.16/million output tokens. At your 16 tokens per second, 1 million tokens is 17 continuous hours of output. Openrouter will charge you 16 cents for that. I think you may want to reevaluate which is the real budget choice here Edit: elaborating, that extra 16GB ram on the Mac to hold the Qwen model costs $400, or equivalently 1770 days of continuous output. All assuming electricity is free
- Karupan 2y agoIt's a no brainer for me cause I already own the MacBook and I don't mind waiting a few extra seconds. Also, I didn't buy the mac for this purpose, it's just my daily device. So yes, I'm sure OpenRouter is cheaper, but I just don't have to think about using it as long as the open models are reasonable good for my use. Of course your needs may be quite different.
- trhway 2y ago>enthusiast AI dev segment i think it isn't about enthusiast. To me it looks like Huang/NVDA is pushing further a small revolution using the opening provided by the AI wave - up until now the GPU was add-on to the general computing core onto which that computing core offloaded some computing. With AI that offloaded computing becomes de-facto the main computing and Huang/NVDA is turning tables by making the CPU is just a small add-on on the GPU, with some general computing offloaded to that CPU. The CPU being located that "close" and with unified memory - that would stimulate development of parallelization for a lot of general computing so that it would be executed on GPU, very fast that way, instead of on the CPU. For example classic of enterprise computing - databases, the SQL ones - a lot, if not, with some work, everything, in these databases can be executed on GPU with a significant performance gain vs. CPU. Why it isn't happening today? Load/unload onto GPU eats into performance, complexity of having only some operations offloaded to GPU is very high in dev effort, etc. Streamlined development on a platform with unified memory will change it. That way Huang/NVDA may pull out rug from under the CPU-first platforms like AMD/INTC and would own both - new AI computing as well as significant share of the classic enterprise one.
- tatersolid 2y ago> these databases can be executed on GPU with a significant performance gain vs. CPU No, they can’t. GPU databases are niche products with severe limitations. GPUs are fast at massively parallel math problems, they anren’t useful for all tasks.
- trhway 2y ago>GPU databases are niche products with severe limitations. today. For the reasons like i mentioned. >GPUs are fast at massively parallel math problems, they anren’t useful for all tasks. GPU are fast at massively parallel tasks. Their memory bandwidth is 10x of that of the CPU for example. So, typical database operations, massively parallel in nature like join or filter, would run about that faster. Majority of computing can be parallelized and thus benefit from being executed on GPU (with unified memory of the practically usable for enterprise sizes like 128GB).
- llm_trw 2y agoFrom the people I talk to the enthusiast market is nvidia 4090/3090 saturated because people want to do their fine tunes also porn on their off time. The Venn diagram of users who post about diffusion models and llms running at home is pretty much a circle.
- dist-epoch 2y agoNot your weights, not your waifu
- Tostino 2y agoYeah, I really don't think the overlap is as much as you imagine. At least in /r/localllama and the discord servers I frequent, the vast majority of users are interested in one or the other primarily, and may just dabble with other things. Obviously this is just my observations...I could be totally misreading things.
- csomar 2y agoAm I the only one disappointed by these? They cost roughly half the price of a macbook pro and offer hmm.. half the capacity in RAM. Sure speed matters in AI, but what do I do with speed when I can't load a 70b model. On the other hand, with a $5000 macbook pro, I can easily load a 70b model and have a "full" macbook pro as a plus. I am not sure I fully understand the value of these cards for someone that want to run personal AI models.
- blurbleblurble 2y agoThen buy two and stack them! Also I'm unfamiliar with macs is there really a MacBook pro with 256GB of RAM?
- csomar 2y agoNo, macbooks pro cap at 128GB. But, still, they are a laptop. It'll be interesting to see if Apple can offer a good counter for the desktop. The mac pro can go to 192Gb which is closer to the 128Gb Digits + your Desktop machine. At $9299 price tag, it's not too competitive but close.
- lr1970 2y ago> It'll be interesting to see if Apple can offer a good counter for the desktop. Mac Pro [0] is a desktop with M2 Ultra and up to 192GB of unified memory. [0] https://www.apple.com/mac-pro/ https://www.apple.com/mac-pro/
- rictic 2y agoHm? They have 128GB of RAM. Macbook Pros cap out at 128GB as well. Will be interesting to see how a Project Digits machine performs in terms of inference speed.
- gnabgib 2y agoAre you, perhaps, commenting on the wrong thread? Project Digits is a $3k 128GB computer.. the best your your $5K MBP can have for ram is.. 128GB.
- behringer 2y agoNot only that, but it should help free up the gpus for the gamers.
- bloomingkales 2y agoJensen did say in recent interview, paraphrasing, “they are trying to kill my company”. Those Macs with unified memory is a threat he is immediately addressing. Jensen is a wartime ceo from the looks of it, he’s not joking. No wonder AMD is staying out of the high end space, since NVIDIA is going head on with Apple (and AMD is not in the business of competing with Apple).
- hkgjjgjfjfjfjf 2y agoYou missed the Ryzen hx ai pro 395 product announcement
- T-A 2y agoFrom https://www.tomshardware.com/pc-components/cpus/amds-beastly-strix-halo-ryzen-ai-max-debuts-with-radical-new-memory-tech-to-feed-rdna-3-5-graphics-and-zen-5-cpu-cores https://www.tomshardware.com/pc-components/cpus/amds-beastly... The fire-breathing 120W Zen 5-powered flagship Ryzen AI Max+ 395 comes packing 16 CPU cores and 32 threads paired with 40 RDNA 3.5 (Radeon 8060S) integrated graphics cores (CUs), but perhaps more importantly, it supports up to 128GB of memory that is shared among the CPU, GPU, and XDNA 2 NPU AI engines. The memory can also be carved up to a distinct pool dedicated to the GPU only, thus delivering an astounding 256 GB/s of memory throughput that unlocks incredible performance in memory capacity-constrained AI workloads (details below). AMD says this delivers groundbreaking capabilities for thin-and-light laptops and mini workstations, particularly in AI workloads. The company also shared plenty of gaming and content creation benchmarks. [...] AMD also shared some rather impressive results showing a Llama 70B Nemotron LLM AI model running on both the Ryzen AI Max+ 395 with 128GB of total system RAM (32GB for the CPU, 96GB allocated to the GPU) and a desktop Nvidia GeForce RTX 4090 with 24GB of VRAM (details of the setups in the slide below). AMD says the AI Max+ 395 delivers up to 2.2X the tokens/second performance of the desktop RTX 4090 card, but the company didn’t share time-to-first-token benchmarks. Perhaps more importantly, AMD claims to do this at an 87% lower TDP than the 450W RTX 4090, with the AI Max+ running at a mere 55W. That implies that systems built on this platform will have exceptional power efficiency metrics in AI workloads.
- 2y ago
- sheepscreek 2y agoThe developers they are referring to aren’t just enthusiasts; they are also developers who were purchasing SuperMicro and Lambda PCs to develop models for their employers. Many enterprises will buy these for local development because it frees up the highly expensive enterprise-level chip for commercial use. This is a genius move. I am more baffled by the insane form factor that can pack this much power inside a Mac Mini-esque body. For just $6000, two of these can run 400B+ models locally. That is absolutely bonkers. Imagine running ChatGPT on your desktop. You couldn’t dream about this stuff even 1 year ago. What a time to be alive!
- stogot 2y agoHow does it run 400B models across two? I didn’t see that in the article
- tempay 2y ago> Nvidia says that two Project Digits machines can be linked together to run up to 405-billion-parameter models, if a job calls for it. Project Digits can deliver a standalone experience, as alluded to earlier, or connect to a primary Windows or Mac PC.
- FuriouslyAdrift 2y agoPoint to point ConnectX connection (RDMA with GPUDirect)
- sliken 2y agoNot sure exactly, but they mentioned linking to together with ConnectX, which could be ethernet or IB. No idea on the speed though.
- HarHarVeryFunny 2y agoThe 1 PetaFLOP spec and 200GB model capacity specs are for FP4 (4-bit floating point), which means inference not training/development. It's still be a decent personal development machine, but not for that size of model.
- 2y ago
- informal007 2y agoI would like to have Mac as my personal computer and digits as service to run llm.
- axegon_ 2y ago> they seem to be doing everything right in the last few years About that... Not like there isn't a lot to be desired from the linux drivers: I'm running a K80 and M40 in a workstation at home and the thought of having to ever touch the drivers, now that the system is operational, terrifies me. It is by far the biggest "don't fix it if it ain't broke" thing in my life.
- mycall 2y agoBuy a second system which you can touch?
- axegon_ 2y agoThat IS the second system (my AI home rig). I've given up on Nvidia for using it on my main computer because of their horrid drivers. I switched to Intel ARC about a month ago and I love it. The only downside is that I have a xeon on my main computer and Intel never really bothered to make ARC compatible with xeons so I had to hack my bios around, hoping I don't mess everything up. Luckily for me, it all went well so now I'm probably one of a dozen or so people worldwide to be running xeons + arc on linux. That said, the fact that I don't have to deal with nvidia's wretched linux drivers does bring a smile to my face.
- sliken 2y agoUse a filesystem that snapshots AND do a complete backup.
- rbanffy 2y agoThis is something every company should make sure they have: an onboarding path. Xeon Phi failed for a number of reasons, but one where it didn't need to fail was availability of software optimised for it. Now we have Xeons and EPYCs, and MI300C's with lots of efficient cores, but we could have been writing software tailored for those for 10 years now. Extracting performance from them would be a solved problem at this point. The same applies for Itanium - the very first thing Intel should have made sure it had was good Linux support. They could have it before the first silicon was released. Itaium was well supported for a while, but it's long dead by now. Similarly, Sun has failed with SPARC, which also didn't have an easy onboarding path after they gave up on workstations. They did some things right: OpenSolaris ensured the OS remained relevant (still is, even if a bit niche), and looking the other way for x86 Solaris helps people to learn and train on it. Oracle cloud could, at least, offer it on cloud instances. Would be nice. Now we see IBM doing the same - there is no reasonable entry level POWER machine that can compete in performance with a workstation-class x86. There is a small half-rack machine that can be mounted on a deskside case, and that's it. I don't know of any company that's planning to deploy new systems on AIX (much less IBMi, which is also POWER), or even for Linux on POWER, because it's just too easy to build it on other, competing platforms. You can get AIX, IBMi and even IBMz cloud instances from IBM cloud, but it's not easy (and I never found a "from-zero-to-ssh-or-5250-or-3270" tutorial for them). I wonder if it's even possible. You can get Linux on Z instances, but there doesn't seem to be a way to get Linux on POWER. At least not from them (several HPC research labs still offer those).
- nimish 2y ago1000% all these ai hardware companies will fail if they don't have this. You must have a cheap way to experiment and develop. Even if you want to only sell a $30000 datacenter card you still need a very low cost way to play. Sad to see big companies like intel and amd don't understand this but they've never come to terms with the fact that software killed the hardware star
- rbanffy 2y ago> Sad to see big companies like intel and amd don't understand this And it's not like they were never bitten (Intel has) by this before.
- numba888 2y ago> I sure wished I held some Nvidia stocks, they seem to be doing everything right in the last few years! They propelled on unexpected LLM boom. But plan 'A' was robotics in which NVidia invested a lot for decades. I think their time is about to come, with Tesla's humanoids for 20-30k and Chinese already selling for $16k.
- GaryNumanVevo 2y agoI bet $100k on NVIDIA stocks ~7 years ago, just recently closed out a bunch of them
- technofiend 2y agoWill there really be a mac mini wirh Max or Ultra CPUs? This feels like somewhat of an overlap with the Mac Studio.
- adolph 2y agoThere will undoubtably be a Mac Studio (and Mac Pro?) bump to M4 at some point. Benchmarks [0] reflect how memory bandwidth and core count [1] compare to processor improvements. Granted, ymmv to your workload. 0. https://www.macstadium.com/blog/m4-mac-mini-review https://www.macstadium.com/blog/m4-mac-mini-review 1. https://www.apple.com/mac/compare/?modelList=Mac-mini-M4,Mac-studio-2023,MacPro-m2-ultra https://www.apple.com/mac/compare/?modelList=Mac-mini-M4,Mac...
- tarsinge 2y ago> I sure wished I held some Nvidia stocks I’m so tired of this recent obsession with the stock market. Now that retail is deeply invested it is tainting everything, like here on a technology forum. I don’t remember people mentioning Apple stock every time Steve Jobs made an announcement in the past decades. Nowadays it seems everyone is invested in Nvidia and just want the stock to go up, and every product announcement is a mean to that end. I really hope we get a crash so that we can get back to a more sane relation with companies and their products.
- wslh 2y agoThe nVidia price is closer (USD 3k) to a top Mac mini but I trust Apple more for the end-to-end support from hardware to apps than nVidia. Not an Apple fanboy but an user/dev, and I don't think we realize what Apple really achieved, industrially speaking. The M1 was launched in late 2020.
- croes 2y agoDid they say anything about power consumption? Apple M chips are pretty efficient.