13 ms·
Only people who never interacted with data center reliability think it's doable to maintain servers with no human intervention.
by lugao 8mo ago
Only people who never interacted with data center reliability think it's doable to maintain servers with no human intervention.
- angled 8mo agoBut … but what if we had solar-powered AI SREs to fix the solar-powered AI satellites… /in space/?
- lugao 8mo agoMaintaining modern accelerators requires frequent hands-on intervention -- replacing hardware, reseating chips, and checking cable integrity. Because these platforms are experimental and rapidly evolving, they aren't 'space-ready.' Space-grade hardware must be 'rad-hardened' and proven over years of testing. By the time an accelerator is reliable enough for orbit, it’s several generations obsolete, making it nearly impossible to compete or turn a profit against ground-based clusters.
- trothamel 8mo agoOn the other hand, Tesla vehicles have similar hardware built into them, and don't require such hands-on intervention. (And that's the hardware that will be going up.)
- lugao 8mo agoCar-grade inference hardware is fundamentally different from data center-grade inference hardware, let alone the specialized, interconnected hardware used for training (like NVLink or complex optical fabrics). These are different beasts in terms of power density, thermal stress, and signaling sensitivity. Beyond that, we don't actually know the failure rate of the Tesla fleet. I’ve never had a personal computer fail from use in my life, but that’s just anecdotal and holds no weight against the law of large numbers. When you operate at the scale of a massive cluster, "one-in-a-million" failures become a daily statistical certainty. Claiming that because you don't personally see cars failing on the side of the road means they require zero intervention actually proves my original point: people who haven't managed data center reliability underestimate the sheer volume of "rare" failures that occur at scale.
- trothamel 8mo agohttps://x.com/elonmusk/status/2017792776415682639 https://x.com/elonmusk/status/2017792776415682639 For what it's worth, this project plans to use Tesla AI5/AI6 hardware for the first launches.
- jonah 8mo agoNot only the sibling comments points, but cars aren't exposed to the radiation of space...
- cloudfudge 8mo agoWell, one car is... and it's a Tesla!
- boutell 8mo agoThank you. The waste heat problem is so bad but no one gets around to mentioning the fact that you can't have AI grade chips and space at the same time.
- elihu 8mo agoDo they need to be maintained? If one compute node breaks, you just turn it off and don't worry about it. You just assume you'll have some amount of unrecoverable errors and build that into the cost/benefit analysis. As long as failures are in line with projections, it's baked in as a cost of doing business. The idea itself may be sound, though that's unrelated to the question of whether Elon Musk can be relied on to be honest with investors about what their real failure projections and cost estimates are and whether it actually makes financial sense to do this now or in the near future.
- lugao 8mo agoAI clusters are heavily interconnected, the blast radius for single component failure is much larger than running single nodes -- you would fragment it beyond recovery to be able to use it meaningfully. I can't get in detail about real numbers but it's not doable with current hardware by a large margin.
- FeepingCreature 8mo agoeh? They're not gonna lay cable in space. The laser links will be retargetable.
- lugao 8mo agoHow are you doing pci express x16 with lasers without fiber optics? Have you touched data center hardware in your life?
- youarentrightjr 8mo agoLasers, space, super geniuses, and most importantly money. You're worrying too much about the details and not enough about the awesomeness. But seriously, why are all the stans in these comments as unknowledgeable as Elon himself? Is that just what is required to stan for this type of garbage?
- 8mo ago
- jmyeet 8mo agoThere are a class of people who may seem smart until they start talking about a subject you know about. Hank Green is a great example of this. For many on HN, Elon buying Twitter was a wake up call because he suddenly started talking about software and servers and data centers and reliability and a ton of people with experience with those things were like "oh... this guy's an idiot". Data centers in space are exactly like this. Your comment (correctly) alludes to this. Companies like Google, Meta, Amazon and Microsoft all have so many servers that parts are failing constantly. They fail so often on large scales that it's expected things like a hard drive will fail while a single job might be running. So all of these companies build systems to detect failures, disable running on that node until it's fixed, alerting someone to what the problem is and then bringing the node back online once the problem it's addressed. Everything will fail. Hard drives, RAM, CPUs, GPUs, SSDs, power supplies, fans, NICs, cables, etc. So all data centers will have a number of technicians who are constantly fixing problems. IIRC Google's ratio tended to be about 10,000 servers per technician. Good technicians could handle higher ratios. When a node goes offline it's not clear why. Techs would take known good parts and basically replacce all of them and then figure out what the problem is later, dispose of any bad parts and put tested good parts into the pool of known good parts for a later incident. Data centers in space lose all of this ability. So if you have a large number of orbital servers, they're going to be failing constantly with no ability to fix them. You can really only deorbit them and replace them and that gets real expensive. Electronics and chips on satellites also aren't consumer grade. They're not even enterprise grade. They're orders of magnitude more reliable than that because they have to deal with error correction terrestial components don't due to cosmic rays and the solar wind. That's why they're a fraction of the power of something you can buy from Amazon but they cost 1000x as much. Because they need to last years and not fail, something no home computer or data center server has to deal with. Put it this way, a hardened satellite or probe CPU is like paying $1 million for a Raspberry Pi. And anybody who has dealt with data centers knows this.
- everfrustrated 8mo agoMight be why he's also investing in building their own fabs - if he can keep the silicon costs low then that flips a lot of the math here.
- 8mo ago
- keepamovin 8mo agoWhoa there, space-faring sysadmin. You really want that off-world contract tho?
- lugao 8mo agoHaha, hard pass on the job. I prefer my oxygen at 1 atm. I'm not a data center technician myself, but I have deep respect for those folks and the complexity they manage. It's quite surprising the market still buys Musk's claims day after day.
- SecretDreams 8mo ago> It's quite surprising the market still buys Musk's claims day after day. More disturbing than surprising.
- andrewinardeer 8mo agoThis guy invented reusable rockets that land themselves. I'm sure xAI is not just one guy. Plenty of talented people work there.
- mrweasel 8mo agoMicrosoft did do the experiment (Project Natick) where they had "datacenters" in pods under the sea with really good results. The idea was simply to ship enough extra capacity, but due to the environment, the failure rates where 1/8th of normal. Still, dropping a pod into the sea makes more sense than launching it into space. At least cooling, power, connectivity and eventual maintenance is simpler. The whole thing makes no sense and is seems like it's just Musk doing financial manipulation again. https://news.microsoft.com/source/features/sustainability/project-natick-underwater-datacenter/ https://news.microsoft.com/source/features/sustainability/pr...
- zarzavat 8mo ago> The whole thing makes no sense and is seems like it's just Musk doing financial manipulation again. It's a fig leaf for getting two IPOs in one. There's no sense in analyzing it any further.
- ryandvm 8mo agoExactly. He can croon about DOGE all day, but the reality is his entire fortune was built on feeding at the trough of government largess. That's why he talks about Mars all the time. He's not stupid enough to think we could actually live there, but damn if he couldn't make a couple trillion skimming off the top of the world's most expensive space program.
- alextingle 8mo agoNo, I think he is that stupid.
- AlexC04 8mo agoRight, let's not forget that he's selling it to himself in an all stock deal. He could have priced it at eleventy kajillion dollars and it would have had the same meaning. He's basically trading two cypto coins with himself and sending out a press release.
- moontear 8mo ago
- donny2018 8mo agoI'd assume datacenters built for space would have different reliability standards. I mean, if a communication satellite (which already has a lot of electronic and computing components) can work unattended, then a satellite working as a server could too.
- vagab0nd 8mo agoYou are right. But in the future we'll be refueling the satellites anyway. Might as well maintain the servers using robots all in one go.
- SilverElfin 8mo agoRight now that’s not the case. Satellites just store whatever fuel they need for orbital adjustments and by default, they fall back to earth and burn up at the end of their life. All the Starlink satellites are configured to fall back to earth within 5 years (the fuel is used to re-raise their orbit). The new proposed datacenters would sit in a higher orbit to avoid debris, allegedly, but that means it is even more expensive to get to them and refuel them, and the potential for future debris is far worse (since it wouldn’t fall back to earth and burn up for centuries or millennia).
- lugao 7mo agoI did some more reading and want to walk back my skepticism a bit. There is actually serious effort going into this, such as Google’s research on space-based AI infrastructure: https://research.google/blog/exploring-a-space-based-scalable-ai-infrastructure-system-design/ https://research.google/blog/exploring-a-space-based-scalabl... They highlight the exact reliability constraint I was thinking of: that replacing failed TPUs is trivial on Earth but impossible in space. Their solution is redundant provisioning, which moves the problem from "operationally impossible" to "extremely expensive." You would effectively need custom, super-redundant motherboards designed to bypass dead chips rather than replace them. The paper also tackles the interconnect problem using specialized optics to sustain high bitrates, which is fascinating but seems incredibly difficult to pull off given that the constellation topology changes constantly. It might be possible, but the resulting hardware would look nothing like a regular datacenter. Also this would require lots of satelites to rival a regular DC which is also very hard to justify. Let's see what the promised 2027 tests will reveal.