6 ms·
> biggest GPU compute cluster in the world right now This is wildly untrue, and most in industry know that. Unfortunately you won't have a source just like I w
by codemac 2y ago
> biggest GPU compute cluster in the world right now
This is wildly untrue, and most in industry know that. Unfortunately you won't have a source just like I won't, but just wanted to voice that you're way off here.
- freedomben 2y ago> This is wildly untrue, and most in industry know that. Unfortunately you won't have a source just like I won't, but just wanted to voice that you're way off here. Sure, we probably can't know for sure who has the biggest as they try to keep that under wraps for competition purposes, but it's definitely not "wildly untrue." A simple search will show that they have if not the biggest, damn near one of the biggest. Just a quick sample: https://nvidianews.nvidia.com/news/spectrum-x-ethernet-networking-xai-colossus https://nvidianews.nvidia.com/news/spectrum-x-ethernet-netwo... https://www.yahoo.com/tech/worlds-fastest-supercomputer-please-stand-145715759.html https://www.yahoo.com/tech/worlds-fastest-supercomputer-plea... https://www.tomshardware.com/pc-components/gpus/elon-musk-took-19-days-to-set-up-100-000-nvidia-h200-gpus-process-normally-takes-4-years https://www.tomshardware.com/pc-components/gpus/elon-musk-to... https://www.capacitymedia.com/article/musks-xais-colossus-cluster-set-for-one-million-gpu-supercomputer-expansion https://www.capacitymedia.com/article/musks-xais-colossus-cl...
- threeseed 2y agoTechnically, it maybe the world's biggest single AI supercomputer. But it ignores Amazon, Google and Microsoft/OpenAI being able to run training workloads across their entire clouds.
- codemac 2y agoi've physically visited a larger one, it is not even a well kept secret. we all see each other at the same airports and hotels.
- nightowl_games 2y agoBecause they are all located in the same small town?
- rvz 2y agoIt is true. [0] [0] https://nvidianews.nvidia.com/news/spectrum-x-ethernet-networking-xai-colossus https://nvidianews.nvidia.com/news/spectrum-x-ethernet-netwo...
- sigh_again 2y agoEven just Meta dwarfs Twitter's cluster, with an estimated 350k H100s by now.
- BrickFingers 2y ago2 months ago Jensen Huang did an interview where he said xAi built the fastest cluster with 100k GPUs.he said "what they achieved is singular, never been done before" https://youtu.be/bUrCR4jQQg8?si=i0MpcIawMVHmHS2e https://youtu.be/bUrCR4jQQg8?si=i0MpcIawMVHmHS2e Meta said they would expand their infrastructure to include 350k GPUs by the end of this year. But, my guess is they meant a collection of AI clusters not a singular large cluster. In the post where they mentioned this, they shared details on 2 clusters with 24k GPUs each.https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure/ https://engineering.fb.com/2024/03/12/data-center-engineerin...
- sigh_again 2y agoWhat's singular is putting 100k H100s in a single machine. Which, yay, cool supercomputer, but the distributed supercomputer with 5 times the machines runs just as fast anyways. Huang is still a CEO trying to prop up his product. He'd tell you putting an RTX4090 in your bathroom to drive an LED screen mirror is unprecedented if it meant it got him more sales and more clout.
- verdverm 2y agoReally? Meta looks to be running larger clusters of Nvidia GPUs already https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure/ https://engineering.fb.com/2024/03/12/data-center-engineerin... This doesn't account for inhouse silicon like Google where the comparison becomes less direct (different devices, multiple subgroups like DeepMind)
- boringg 2y agoI don't think you've been paying attention to the industry even though your posturing like an insider.
- enslavedrobot 2y agoThe distinction is that larger installations cannot form a single network. Before xAI's new network architecture, only around 30k GPUs could train a model simultaneously. It's not clear how many can train together with xAI's new approach, but apparently it is >100k.