6 ms·
Cloud TPU v5p and AI Hypercomputer
- StephenSmith 3y agoThis is the bigger headline than their Gemini release. AI is all about how much compute dollars it can generate for the cloud providers. Google is trying to make sure Microsoft doesn't monopolize AI compute.
- verdverm 3y agoGiven the TPUv5 improves perf/$$$, it would seem to be at odds with your comment. I can now get more done with the same spend. Kelsey Hightower told me at a GopherCon (many years ago) that Google doesn't run any internal workloads on third-party GPUs mainly because it costs significantly more (b/c cooling iirc), though they are happy to help you run your workloads on such GPUs.
- webmaven 3y agoWhich (potentially) draws compute spend from Microsoft Azure to Google Cloud.
- verdverm 3y agoI should have quoted the generalized middle statement I was responding to > AI is all about how much compute dollars it can generate for the cloud providers. If the providers wanted to extract more money they would not create custom hardware which reduces overall costs and prices to users. I would argue that this is actually more about ensuring NVidia doesn't have a monopoly on hardware and alleviates us from having to pay for Nvidia profits through our cloud providers. Azure is going down the same path here: https://www.theverge.com/2023/11/15/23960345/microsoft-cpu-gpu-ai-chips-azure-maia-cobalt-specifications-cloud-infrastructure https://www.theverge.com/2023/11/15/23960345/microsoft-cpu-g...
- qeternity 3y ago> If the providers wanted to extract more money they would not create custom hardware which reduces overall costs and prices to users. Extracting money is about margins, not revenues. If they reduce your costs (and their revenue) by 20% with a TPU, but they can produce TPUs for 50% less than buying gear from Nvidia, it's still a profitable move.
- verdverm 3y agoExactly, if I end up paying less and the cloud also makes more money, seems like a win for everyone The "extracting" word typically comes with abusive connotations when used in the context of money, which doesn't feel like the right word for the win-win outcomes imho
- dekhn 3y agoGoogle runs internal workloads on third-party GPUs (especially as TPUs weren't very good at sparse for a long time). Hightower was simply wrong.
- holografix 3y agoHightower’s biggest skill is sounding very knowledgeable and as if he has the keys to a gate of useful secrets.
- verdverm 3y agoKelsey is kind, thoughtful, and never puts another person down, I would say empathy is his best skill. He also shares what he has learned for free, rather than putting paywalls in front of his content, which is quite rare these days.
- freen 3y agoJevons paradox. Higher efficiency results in greater utilization.
- DeathArrow 3y agoHow does it compare to Nvidia A100?
- ZeroCool2u 3y agoI'm sure they've done benchmarking comparisons with the A100 and H100, but since they're not sharing them it seems unlikely they show the TPU's in a favorable light. Plus, they pick up more spend every time someone runs these benchmarks for themselves.
- deleted 3y ago[deleted]
- RicoElectrico 3y agoThe most important reason all the Big Tech companies spin their own AI chips is that NVIDIA margins are so insane. A purchase from NVIDIA at the hyperscaler volume is on the same order of magnitude as spinning your own chip.
- terafo 3y agov4 chips are basically the same in bf16 performance as A100(slower in int8).
- modeless 3y agoGenerally TPUs don't beat Nvidia per chip. The real reason Google likes them is they are much cheaper in TCO for Google. Of course they may not decide to pass on all the savings to you.
- ithkuil 3y agoI wonder why they don't try to compete as much as they can.
- andromeduck 3y agoThey do - it's just a different design point.
- leetharris 3y agoI've been working with people at GCP for months to get the right provisioning for TPUs for my company that spends many millions per year on GPU compute. They take weeks to respond to anything, they change their minds constantly, you can never trust anything anyone says, their internal communication is a complete disaster, and someone recently told me they outsource a lot of their GCP personnel. We went back to AWS and had a whole fleet of GPUs up and running within the week. This is my 3rd extremely bad experience on Google Cloud. My last unicorn startup had several GCP-caused P0 production issues. They would update something internally with no announcement to customers and our production workloads would completely break out of the blue. It would usually take them days to weeks to fix it even with us spending tens of millions with them and calling support constantly. Everyone at our company was baffled at how bad the experience was compared to every other cloud provider. I would not put anything serious there and I would never partner with GCP again.
- shiftpgdn 3y agoYou must not have been spending million per year. A friend's company spends 10 million/year with GCP, which isn't huge, and can have an engineer from any group in a meeting the next day after a high priority issue. How frequently are you engaging your account reps? You should be able to get the ear of a PM within 48 hours in most cases.
- cornel_io 3y agoYeah, I spend a small fraction of that and have found GCP support through our account rep to be extremely good, at least on par with what AWS provides. Maybe the particular rep makes a big difference?
- Palmik 3y agoThe parent said "my company that spends many millions per year".
- shiftpgdn 3y agoI understand that and am calling that claim into question.
- lhl 3y agoRecently I've been using GCP to train a model, some notes: * Like @leetharris, credits were pulled/not distributed, what was promised had to be cajoled out with most of it being sent to some weird SaaS product that we'll never use * The GCP rep literally ghosted us halfway through the month where it when we had some expiring credits and were in the middle of training * Not that the credits mattered, our quota requests for lifting GPU or TPU was rejected twice. It was impossible to get any GPUs that were within our credits, even writing a script to try to look for machines for weeks didn't work. * Right after the credits expired, suddenly our last quota request, which was hanging around for weeks was approved. I assume they have an internal system setup to do that, but like we literally couldn't pay for GCP if we wanted to. * Also, GCP rates are like 2-4X the market rate. Like you can get an H100-80 from Runpod (and actually get one) for what GCP charges for an A100-40. Basically, the lesson learned was that no one should ever depend on GCP unless your time is worthless and you're not serious about getting any work done. They can go suck eggs.
- mg 3y agoDoes Google rely on TSMC to build the TPU chips?
- HarHarVeryFunny 3y agoMaybe indirectly. I was surprised to learn (just Googled it!) that Google TPU chips are mostly designed by Broadcom, who (being mostly fabless themselves) do in turn use a variety of companies such TSMC, Global Foundaries, etc to make them. Not sure if TPUs use latest cutting edge nodes - if so then presumably it is specifically TSMC. https://www.theinformation.com/articles/to-reduce-ai-costs-google-wants-to-ditch-broadcom-as-its-tpu-server-chip-supplier https://www.theinformation.com/articles/to-reduce-ai-costs-g... https://en.wikipedia.org/wiki/Broadcom_Corporation https://en.wikipedia.org/wiki/Broadcom_Corporation
- pclmulqdq 3y agoStill no FP8 from Google. Surprising, given how effective it seems to be for both training and inference. Although it's not that surprising given that the primary customer of TPUs is Google itself, and they tend to stick themselves on weird little tech islands.
- buildbot 3y agoEven below FP8 works: https://arxiv.org/abs/2310.10537 https://arxiv.org/abs/2310.10537 (6 bit and lower)
- pclmulqdq 3y agoI have seen FP4 as a proposed format, as well as batch floating point with FP8 numbers (batch floating point means n mantissas for every exponent - an old DSP trick), resulting in ~4-5 bits per number. I'm just disappointed that Google isn't taking quantization very seriously. Edit: As the commenter below points out, "block floating point" is the common name, not "batch floating point."
- buildbot 3y agoThe Microscaling paper and the following MX OCP spec is along the lines of what you call batch floating point, though I believe the original 1963(?) work on it called it "blocked" floating point. Datatypes are really tricky. Hardware designers tend to be conservative in my experience, and don't want to waste die space on things that might not be useful. Edit - The original: https://www.abebooks.com/first-edition/Rounding-Errors-Algebraic-Processes-Prentice-Hall-Series/31429549335/bd https://www.abebooks.com/first-edition/Rounding-Errors-Algeb...
- camdenlock 3y ago> To request access Google is so fucking lame these days
- faeriechangling 3y agoThey’ve used this scheme since the launch of gmail
- 0cf8612b2e1e 3y agoSerious question, are there open source designs available for RISC today that do “good enough” matrix multiplication? I have no doubt that Nvidia has extensive optimizations to get SOTA performance, but I am curious what is attainable off the shelf. If you could design a 5nm chip, would it be possible to hit 15% of a NVidia chip? Significantly more? Of course, there is more to a GPU than just the matrix multiplication, but I am wondering how much effort it would take to get something off the ground for the well financed organization. Presumably China is actively finding such efforts.
- buildbot 3y agoProbably. There are several companies going with this approach. The real trick is the software and ecosystem around it. It does not matter if you can outperform an H100 if you spend weeks trying to get software to work or debugging if a NaN is a hardware error or your bad code :)
- kaycebasques 3y agoThere were a bunch of presentations about matrix multiplication at the RISC-V summit last month. I'm not sure if any of the presented hardware is open source but maybe those videos are a good lead for tracking some down? https://www.youtube.com/playlist?list=PL85jopFZCnbMfMRR25ENcRkhhAUGwP5C5 https://www.youtube.com/playlist?list=PL85jopFZCnbMfMRR25ENc...
- htrp 3y agoRain is the one that most people (including Altman) are backing.
- modeless 3y ago> large LLM models I'm not usually one to point out redundancies like this but this one seems egregious.
- stabbles 3y agoThis is in contrast to small large language models like GPT3
- ithkuil 3y agoI think the redundant part is "models: Large Large Language Model Models Small Large Language Model Models
- modeless 3y agoThey're both redundant. There's no such thing as a small large model. It's a small model or a medium model or a large model.
- ithkuil 3y agoOnce upon a time there was a family of Bigfoot. Papa Bigfoot, Mommy Bigfoot and Little Bigfoot. Papa Bigfoot was the biggest Bigfoot of them all. Mommy Bigfoot wasn't as big as her husband, but she was still a bigger Bigfoot than her daughter Little Bigfoot, who was the smallest Bigfoot of the family. One day Little Bigfoot slipped in a stream and hurt her foot. The little Little Bigfoot foot hurt so much and she cried a lot
- riku_iki 3y agoI like how they launch hard before end of the year performance review.
- az226 3y agoLol. Without benchmarks against H100 you can’t take this seriously.