5 ms·
I'll note CoreWeave didn't commission the very first Blackwells until ~Jan 2025 so it hasn't been very long for RMAs. Furthermore, for training even when the NV
by kurthr 21d ago
I'll note CoreWeave didn't commission the very first Blackwells until ~Jan 2025 so it hasn't been very long for RMAs. Furthermore, for training even when the NVL backplane can use the remaining GPUs a ~20% drop in performance means it isn't in the training cluster.
I've personally heard this from several sources in the data centers (installers, training, network). It's not uncommon for 10% of racks to fail on delivery. I hear that's improved somewhat from GB200 to GB300, but the number of FW updates from the time they ship, until they're commissioned is >>10. If an HBM or GPU or backplane supply/cooling fails, it is basically not swappable or repairable. You have a "dead" rack, and deliveries are on allocation so you don't get a replacement for months (eg some "RMAs" for early delivered parts in late 2025 are still dead racks 9 months later). "Tray" swaps are technically possible, but still quite rare, perhaps because debugging takes as much time as commissioning a new rack.
I don't want to out anyone, but these are similar comments:
https://www.linkedin.com/posts/neelmaster1_aiinfrastructure-inference-nvl72-activity-7492641179134750721--lpJ https://www.linkedin.com/posts/neelmaster1_aiinfrastructure-...
https://www.hostzealot.com/blog/news/nvidia-gb200-nvl72-is-not-yet-ready-for-training-advanced-ai-models https://www.hostzealot.com/blog/news/nvidia-gb200-nvl72-is-n...
https://introl.com/blog/gb200-nvl72-deployment-72-gpu-liquid-cooled https://introl.com/blog/gb200-nvl72-deployment-72-gpu-liquid...