9 ms·
Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
Hey folks! We're Alex and Evan, and we're working on putting together a 512 H100 compute cluster for startups and researchers to train large generative models on.
- it runs at the lowest possible margins (<$2.00/hr per H100)
- designed for bursty training runs, so you can take say 128 H100s for a week
- you don’t need to commit to multiple years of compute or pay for a year upfront
Big labs like OpenAI and Deepmind have big clusters that support this kind of bursty allocation for their researchers, but startups so far have had to get very small clusters on very long term contracts, wait months of lead time, and try to keep them busy all the time.
Our goal is to make it about 10-20x cheaper to do an AI startup than it is right now. Stable Diffusion only costs about $100k to train -- in theory every YC company could get up to that scale. It's just that no cloud provider in the world will give you $100k of compute for just a couple weeks, so startups have to raise 20x that much to buy a whole year of compute.
Once the cluster is online, we're going to be pretty much the only option for startups to do big training runs like that on.
- ucarion 3y agoWishing y'all the best of luck. This would be huge for a lot of folks.
- sashank_1509 3y agoCorrect me if I’m wrong but doesn’t Lambda Labs already provide them at 1.89$? What’s the point if you’re starting out not the cheapest
- davidmurphy 3y agoLooks like their site is quoting a rate of $1.99 now https://lambdalabs.com/ https://lambdalabs.com/
- version_five 3y agoSee this post above: https://news.ycombinator.com/item?id=36935032 https://news.ycombinator.com/item?id=36935032 Price and market depth are very different things
- agajews 3y agoAh that’s only if you pay for 3 years of compute upfront. Most startups, especially the small ones, really can’t afford that
- AndrewKemendo 3y agoThe billion dollar question is: Who is funding this? Cause if it’s VC then it’s going to have the same fate as everything else after 5-7 years. I hope y’all have as innovative of a business model. You’ll need it if you want to do what you’re doing now for more than a few years
- fragmede 3y agoWhat's wrong with doing something profitable for a few years? H100's in a couple of years will be like having a cluster of K80's today. Not everything has grow to have the appetite of Galactus and swallow a whole planet. Making single digit millions of dollars over a couple of years is still worthwhile, especially if it helps others and moves humanity forwards. This project isn't ever going to want to try and compete with AWS, so no, it's not a billion dollar question. $20 Million, yeah.
- AndrewKemendo 3y agoHey I agree! That’s why I’m asking because a “bootstrapped” company like you describe has a future… One backed by VC doesn’t I mean they may have a future but not like you describe
- constantly 3y agoYou’re completely right in everything you say about growing sustainably and making some money over time. But if this project is VC that all goes out the window and it won’t be profitable unless it massively galactus scales to compete with AWS in 5-7 years, and will fail after that almost certainly, like the vast, vast majority of VC projects.
- agajews 3y ago[dead]
- sillysaurusx 3y agoI hope you succeed. TPU research cloud (TRC) tried this in 2019. It was how I got my start. In 2023 you can barely get a single TPU for more than an hour. Back then you could get literally hundreds, with an s. I believed in TRC. I thought they’d solve it by scaling, and building a whole continent of TPUs. But in the end, TPU time was cut short in favor of internal researchers — some researchers being more equal than others. And how could it be any other way? If I made a proposal today to get these H100s to train GPT to play chess, people would laugh. The world is different now. Your project has a youthful optimism that I hope you won’t lose as you go. And in fact it might be the way to win in the long run. So whenever someone comes knocking, begging for a tiny slice of your H100s for their harebrained idea, I hope you’ll humor them. It’s the only reason I was able to become anybody.
- LoganDark 3y ago> In 2023 you can barely get a single TPU for more than an hour. Um. Can't you order them from coral.ai and put them in an NVMe slot? Or are the cloud TPUs more powerful?
- whimsicalism 3y agoTPU pod is not sold by google, edge tpu is different
- LoganDark 3y agoSo the cloud TPUs are more powerful...? Or what are you saying?
- whimsicalism 3y agoyes
- sillysaurusx 3y agoYeah, it’s a silly branding thing. One TPU (not even a pod, just a regular old TPUv2) has 96 CPU cores with 1.4TB of RAM, and that’s not even counting their hardware acceleration. I’d love to buy one.
- williamstein 3y agoHow does this compare to https://lambdalabs.com/ https://lambdalabs.com/ ?
- jorlow 3y agoYou can usually only get a few h100s at a time unless you're committed to reserved instances (for a longer time period)
- wongarsu 3y agoVery similar price, but from what I gather very different model. One important difference might be if you regularly run short-ish training runs over many GPUs. Lambdalabs might not have 256 instances to give you right now. With OP you are basically buying the right to put jobs in the job queue for their 512 GPU cluster, so running a job that needs 256 GPUs isn't an issue (though you might wait behind someone running a 512 GPU job). No idea how capacity at lambdalabs actually looks like though. Does anyone have insight how easy it is to spin up more than 2-3 instances up there?
- agajews 3y agoYeah it’s pretty hard to find a big block of GPUs that you can use for a short time, esp if you need infiniband for multinode training. Lambda I think needs a min reservation of 6-12 months if you want IB.
- flaque 3y agoAh, we're running a medium amount of compute at zero-margin. The point is not to go sell the Fortune 500, but to make sure a grad student can spend a $50k grant. Right now, it's pretty easy to get a few A/H100s (Lambda is great for this), but very hard to get more than 24 at a reasonable price ($~2 an hour). One often needs to put up a 6+ month commitment, even when they may only want to run their H100s for an 8 hour training run. It's the right business decision for GPU brokers to do long term reservations and so on, and we might do so too if we were in their shoes. But we're not in their shoes and have a very different goal: arm the rebels! Let someone who isn't BigCorp train a model!
- latchkey 3y ago554 5.7.1 <evan@sfcompute.org>: Relay access denied 554 5.7.1 <alex@sfcompute.org>: Relay access denied
- flaque 3y ago!!!!!! fixing this. For the moment, evan at roomservice dot dev
- deleted 3y ago[deleted]
- fragmede 3y agofwiw, https://roomservice.dev/ https://roomservice.dev/ is currently a 404
- flaque 3y agoAh yeah, that's normal! Was from my old CRDT company, and works as a good emergency email while we debug our DNS.
- deleted 3y ago[deleted]
- fragmede 3y agoI assume it was a Take3 reference. I wanted to point it out, in case it was supposed to return more than a 404.
- latchkey 3y agohttp != smtp roomservice.dev. 60 IN MX 5 alt1.aspmx.l.google.com. roomservice.dev. 60 IN MX 5 alt2.aspmx.l.google.com. roomservice.dev. 60 IN MX 1 aspmx.l.google.com. roomservice.dev. 60 IN MX 10 alt3.aspmx.l.google.com. roomservice.dev. 60 IN MX 10 alt4.aspmx.l.google.com. roomservice.dev. 60 IN MX 15 4ig53n4pw7p3cuxm7n7xi7dpuyq6722aipexvhkngzbd2e4mudmq.mx-verification.google.com.
- whimsicalism 3y agoI am super interested in AI on a personal level and have been involved for a number of years. I have never seen a GPU crunch quite like it is right now. To anyone who is interested in hobbyist ML, I highly highly recommend using vast.ai
- williamstein 3y agoMany thanks for posting about vast.ai, which I had never heard of! It's a sort of "gig economy/marketplace" for GPU's. The first machine I tried just now worked fine, had 512GB of RAM, 256 AMC CPUs, an A100 GPU, and I got about 4 minutes for $0.05 (which they provided for free).
- whimsicalism 3y agoThe only caveat is it is not really appropriate for private usecases. Also, many of the available options clearly are recycled crypto mining rigs which have somewhat odd configurations (poor gpu bandwidth, low cpu ram).
- quickthrower2 3y agoDepends on what you class as hobbyist but I am running a T4 for a few minutes to get acquainted with tools and concepts and I found modal.com really good for this. They resell AWS and GCP at the moment. They also have A100 but T4 is all I need for now.
- whimsicalism 3y agoSignificantly more expensive than equivalent 3090 configuration if you can do model parallelism
- quickthrower2 3y agoWhat do you mean by this? I use less than the $30/m free included usage. I am guessing you mean at some point just buy your own 3090 as it will be cheaper than paying a cloud per second for a server-grade Nvidia setup.
- rsync 3y ago"Once the cluster is online ..." Where will the cluster be hosted ? May I suggest that you get your IP transit from he.net ?
- deleted 3y ago[deleted]
- fragmede 3y agoNot to mention, San Francisco is not known for having cheap real estate, nor is it known for having cheap electricity. My last (residential) bill to PGE, I paid $0.50938/KWh at peak.
- vladgur 3y agoWhile business rates may be different, California cannot be a sensible place to host power-hungry infrastructure - our electrical rates are easily 5-8 times of other locations within the US
- moneycantbuy 3y agoHow did you get the money to buy 512 H100s?
- turbobooster 3y ago[dead]
- taminka 3y agoask no questions hear no lies
- rvnx 3y agoEDIT: They seem to be in a raising fund / debt stage. Great initiative
- williamstein 3y agoTheir announcement says "We can probably get a good deal from a bank [...]", so maybe they don't just have 20M USD sitting around.
- rvnx 3y agoWell, this pushes me even further in the direction that they are actually good guys that need support, and that they are trying to bring a good deal on the table :)
- deleted 3y ago[deleted]
- herval 3y agounrelated to this specific initiative, but - I keep seeing a lot of announcements of huge VC rounds around what's effectively datacenters for GPUs. Curious about the math behind that - I feel like those things get obsolete so fast, it's almost like the whole scooter rental thing, where the unit economics doesn't add up. Anyone have an insight?
- 3y ago
- 29athrowaway 3y agoDuring a gold rush, sell shovels. When was the last time you spoke to a chatbot?
- lulunananaluna 3y agoDownvoted by others, yet very true. This is a valid business model, nothing to be ashamed about it.
- netsec_burn 3y agoFor me, today and almost every day since the beginning of this year. Not sure if that saying applies here.
- version_five 3y agoChatbot in the sense I think you mean is a horrible application. Millions of people are using large language models daily though.
- nilsbunger 3y agoI love the idea of community assets. could it be the start of a GPU co-op?
- samstave 3y agoSerious Q, as I dont know Twitters internal infra at all... but with a shrinking in revenue from ads, or maybe less engagement by users, and the influx of Threads - maybe twitter can use from slices of its infra (even if its rack space, VMs, Containers, connectivity, who knows what, to support startups such as this? Basically twitter devolves into the Colos of the late 90s :-) - For those who didnt notice, it was tongue in cheek.
- version_five 3y agoI've generally tried to give Twitter the benefit of the doubt but I would never trust them as an infrastructure provider in their current incarnation. Reliability and consistency have been so far from their focus.
- aionaiodfgnio 3y agoWould you really trust a company that doesn't pay its rent to run your infrastructure?
- mike_d 3y agoGenerally when you just stop paying your bills the datacenter holds your hardware and eventually auctions it off to cover some of your debt. I seriously doubt Twitter has any access to the two of three datacenters Elon decided to not pay for.
- fragmede 3y agoFor consumer-grade cards, that's already here. Make money off your GPU with vast.AI https://cloud.vast.ai/host/setup https://cloud.vast.ai/host/setup
- PartiallyTyped 3y ago
- resonance1994 3y agoJust curious, do you guys use renewable energy to power your cluster?
- kaycebasques 3y agoHi, SF lover [1] here. Anything interesting to note about your name? Will your hardware actually be based in SF? Any plans to start meetups or bring customers together for socializing or anything like that? [1] We have not gone the way of the Xerces blue [2] yet... we still exist! [2] https://en.wikipedia.org/wiki/Xerces_blue https://en.wikipedia.org/wiki/Xerces_blue
- agajews 3y agoAh the hardware isn’t gonna be in SF (not the cheapest datacenter space) But I do think a lot of our customers will be out here —- SF is still probably the best place to do startups. We just have so many more people doing hard technical stuff here. Literally every single place I’ve lived in SF there’s been another startup living upstairs or downstairs Good idea to host some in person events!
- menthe 3y ago> SF is still probably the best place to do startups. now that's a hot take if I ever saw one
- deleted 3y ago[deleted]
- whack 3y ago> Rather than each of K startups individually buying clusters of N gpus, together we buy a cluster with NK gpus... Then we set up a job scheduler to allocate compute In theory, this sounds almost identical to the business model behind AWS, Azure, and other cloud providers. "Instead of everyone buying a fixed amount of hardware for individual use, we'll buy a massive pool of hardware that people can time-share." Outside of cloud providers having to mark up prices to give themselves a net-margin, is there something else they are failing to do, hence creating the need for these projects?
- beachy 3y agoAWS and Azure would slit their own throats before they created a way for their customers to pool instances to save money. They want to do that themselves, and keep the customer relationship and the profits, instead of giving them to a middleman or the customer.
- jiggawatts 3y agoIt’s just corporate profits combined with market forces, not a some sort of malicious conspiracy. You can rent a 2-socket AMD server with 120 available cores and RDMA for something like 50c to $2 per hour. That’s just barely above the cost of the electricity and cooling! What do you want, free compute just handed to you out of the goodness of their hearts? There is incredible demand for high-end GPUs right now, and market prices reflect that.
- mikeravkine 3y agoSorry where are these .50c many core servers you speak of exactly?
- jiggawatts 3y agoAzure's HB120rs_v3 size is about 36c per hour right now with Spot pricing in East US. These use 3rd generation AMD EPYC "Milan" processors. The instances with the 4th generation "Genoa-X" processors (HB176rs_v4) cost about $2.88 per hour. The HX176rs_v4 model with 1.7 TB of memory is $3.46 per hour. https://learn.microsoft.com/en-us/azure/virtual-machines/hbv3-series https://learn.microsoft.com/en-us/azure/virtual-machines/hbv... https://learn.microsoft.com/en-us/azure/virtual-machines/hbv4-series https://learn.microsoft.com/en-us/azure/virtual-machines/hbv... https://learn.microsoft.com/en-us/azure/virtual-machines/hx-series https://learn.microsoft.com/en-us/azure/virtual-machines/hx-...
- bnr4u 3y agoHaving hosted infrastructure in CA at multiple colos. I would advise you to host it elsewhere if you can, cost of power, other infrastructure is much higher in CA than AZ or NV.
- ec109685 3y agoPower seems like a very small amount of cost of compute when it comes to GPU’s.
- version_five 3y agoFWIW I tired to look up some numbers, i found California "industrial" electricity at $0.18/Kwh https://www.eia.gov/electricity/monthly/epm_table_grapher.php?t=epmt_5_6_a https://www.eia.gov/electricity/monthly/epm_table_grapher.ph... and H100s using 300-700w https://www.nvidia.com/en-us/data-center/h100/ https://www.nvidia.com/en-us/data-center/h100/ which implies a worst case marginal cost of .18*.7 = $.126 / gpu / hour. Looks like Montana is cheapest at ~$.05 / kwh which would bring that down to $.035. So there may be about a $0.09 California premium (vs the absolute cheapest possibility), which as you say is a small amount of the total cost, but could be material for large workloads.
- jedberg 3y agoRetail residential power in the city of Santa Clara is $0.15/KwH, I'm sure commercial could be less. Especially if you throw some solar panels on the roof. The most expensive part would be the land, but honestly there is some pretty cheap land outside the cities.
- itissid 3y agoNoob Thought: So this would be a blue print on how a mid tier universities with older large compute cluster ops could do things in 2023 to support large LLM research? Perhaps its also a way for freshly applying grad students to look at a university looking to do research in LLMs that requires scale...
- itissid 3y agoLike to clarify, a new grad students could look at the current group and ask "Hey I know you are working on LLMs, but how many $$ of your grant are dedicated to how many TPU hours per grad student?"
- mackid 3y agoNat Friedman and Daniel Gross setup a 2,512 H100 cluster [1] for their startups, with a very similar “shared” model. Might be interesting to connect with them. [1] https://andromedacluster.com/ https://andromedacluster.com/
- flaque 3y agoNat & Daniel’s cluster is great, and we fully recommend startups seek out this option as well. Nat & Daniel are some of the best investors one can have
- dudus 3y agoI know AWS/GCP/Azure have overhead and I understand why so many companies choose to go bare metal on their ops. I personally rarely think it's worth the time and effort, but I get that with scale saving can be substantial. But for AI training? If the public cloud isn't competitive even for bursty AI training, their margins are much higher than I anticipated. OP mentions 10-20x cost reduction? Compared to what? AWS?
- jerjerjer 3y agoAWS offers p5.48xlarge which is 8xH100 for $98.32, so 12.29$ per hour per H100 - ~6x the price.
- metadat 3y agoWill it be a Slurm cluster, or what kind of scheduler is SFC planning to use?
- rushingcreek 3y agoI love this. Us at Phind.com would love to be a part of this.
- netcraft 3y agoHonest question I don’t know how to consider: are we further along or behind with AI given crypto’s use of GPUs? Has the same cards bought for mining furthered AI, or maybe that demand lead to more research into GPUs and what they can do - or would we be further along if we weren’t wasting these cards on mining?
- fragmede 3y agoEthereum's (thrice delayed) move to PoS put a glut of GPUs on the market, just in time for the AI boom to swallow them back up, so I think it ended up okay. NVDA certainly had a great few days in the market thanks to AI though.
- coffeebeqn 3y agoEth was mined mainly on consumer GPUs which as far as i understand, have too little VRAM for most AI training
- wodenokoto 3y ago> It's just that no cloud provider in the world will give you $100k of compute for just a couple weeks I've never had to buy very large compute, but I thought that was the whole point of the cloud
- jeepers6 3y agoPlease take this question without prejudice. Is it accurate to say you’re willing to go into ~20,000,000 USD debt to sell discounted computer-as-a-service to researchers/startups, but unwilling to go into debt to sponsor the undergraduate degrees of ~100-500 students at top-tier schools? (40k - 200k USD per degree) Or, you know, build and fund a small public school/library or two for ~5 years?
- deleted 3y ago[deleted]
- PettingRabbits 3y agoWhat kind of hardware setup are you planning out? Colocation, roll-your-own data center, something in between? Any thoughts on what servers the GPUs will be housed in?
- orGANicWeb 3y ago[flagged]
- orGANicWeb 3y agoHow are you going to sell access and divide the resources?