8 ms·
This article doesn't mention TPUs anywhere. I don't think it's obvious for people outside of google's ecosystem just how extraordinarily good the JAX + TPU ecos
by thunderbird120 1y ago
This article doesn't mention TPUs anywhere. I don't think it's obvious for people outside of google's ecosystem just how extraordinarily good the JAX + TPU ecosystem is. Google several structural advantages over other major players, but the largest one is that they roll their own compute solution which is actually very mature and competitive. TPUs are extremely good at both training and inference[1] especially at scale. Google's ability to tailor their mature hardware to exactly what they need gives them a massive leg up on competition. AI companies fundamentally have to answer the question "what can you do that no one else can?". Google's hardware advantage provides an actual answer to that question which can't be erased the next time someone drops a new model onto huggingface.
[1]https://blog.google/products/google-cloud/ironwood-tpu-age-of-inference/ https://blog.google/products/google-cloud/ironwood-tpu-age-o...
- noosphr 1y agoAnd yet google's main structural disadvantage is being google. Modern BERT with the extended context has solved natural language web search. I mean it as no exaggeration that _everything_ google does for search is now obsolete. The only reason why google search isn't dead yet is that it takes a while to index all web paged into a vector database. And yet it wasn't google that released the architecture update, it was hugging face as a summer collaboration between a dozen people. Google's version came out in 2018 and languished for a decade because it would destroy their business model. Google is too risk averse to do anything, but completely doomed if they don't cannibalize their cash cow product. Web search is no longer a crown jewel, but plumbing that answering services, like perplexity, need. I don't see google being able to pull off an iPhone moment where they killed the iPod to win the next 20 years.
- visarga 1y ago> Modern BERT with the extended context has solved natural language web search. I mean it as no exaggeration that _everything_ google does for search is now obsolete. The web UI for people using search may be obsolete, but search is hot, all AIs need it, both web and local. It's because models don't have recent information in them and are unable to reliably quote from memory.
- nroets 1y agoAnd models often makes reasoning errors. Many users will want to check that the sources substantiate the conclusion.
- fennecfoxy 1y agoAs they should do even for a Google search. I see search engines as a dripfeed from a firehose, not some magical thing that's going to get me the 100% correct 100% accurate result. Humans are the most prolific liars; I could never trust search results anyway since Google may find something that looks right but the author may be heavily biased, uninformed and all manner of other things anyways.
- vidarh 1y agoThe point is that the secret sauce in Google's search was better retrieval, and the assertion above is that the advantage there is gone. While crawling the web isn't a piece of cake, it's a much smaller moat than retrieval quality was.
- pixl97 1y agoEh, I don't really see that. Crawling the web has a huge moat because a huge number of sites have blocked 'abusive' crawlers except Google and possibly Bing. For example just try to crawl sites like Reddit and see how long before you're blocked and get a "please pay us for our data" message.
- literalAardvark 1y agoMy experience running a few hundred very successful shops (hundreds of thousands of orders per month) is that there's no need for quotes around 'abusive'. 95% of our load is from crawlers, so we have to pick who to serve. If they want our data all they need to do is offer a way for us to send it, we're happy to increase exposure and shopping aggregation site updates are our second highest priority task after price and availability updates.
- 1y ago
- podnami 1y agoDo we have insights on whether they knew that their business model was at risk? My understanding is that OpenAI’s credibility lies in seeing the potential of scaling up a transformer-based model and that Google was caught off guard.
- dash2 1y agoThey can just plug the google.com web page into their AI. They already do that.
- fragmede 1y agobut because users are used to doing that for free, they can't charge money for that, but if they don't charge money for that, and no one's seeing ads, then where does they money come from?
- eitally 1y agoWell, it clearly affects search ads, but in terms of revenue streams Google is already somewhat diversified: 1. Search ads (at risk of disintermediation) 2. Display ads (not going anywhere) 3. Ad-supported YouTube 4. Ad-supported YouTube TV 5. Ad-supported Maps 6. Partnership/Ad supported Travel, YouTube, News, Shopping (and probably several more) 7. Hardware (ChromeOS licensing, Android, Pixel, Nest) 8. Cloud There are probably more ad-supported or ad-enhanced properties, but what's been shifting over the past few years is the focus on subscription-supported products: 1. YouTube TV 2. YouTube Premium 3. GoogleOne (initially for storage, but now also for advanced AI access) 4. Nest Aware 5. Android Play Store 6. Google Fi 7. Workspace (and affiliated products) In terms of search, we're already seeing a renaissance of new options, most of which are AI-powered or enhanced, like basic LLM interfaces (ChatGPT, Gemini, etc), or fundamentally improved products like Perplexity & Kagi. But Google has a broad and deep moat relative to any direct competitors. Its existential risk factors are mostly regulation/legal challenge and specific product competition, but not everything on all fronts all at once.
- petesergeant 1y ago> Google is too risk averse to do anything, but completely doomed if they don't cannibalize their cash cow product. Google's cash-cow product is relevant ads. You can display relevant ads in LLM output or natural language web-search. As long as people are interacting with a Google property, I really don't think it matters what that product is, as long as there are ad views. Also: > Web search is no longer a crown jewel, but plumbing that answering services, like perplexity, need This sounds like a gigantic competitive advantage if you're selling AI-based products. You don't have to give everyone access to the good search via API, just your inhouse AI generator.
- michaelt 1y agoKodak was well placed to profit from the rise of digital imaging - in the late 1970s and early 1980s Kodak labs pioneered colour image sensors, and was producing some of the highest resolution CCDs out there. Bryce Bayer worked for Kodak when he invented and patented the Bayer pattern filter used in essentially every colour image sensor to this day. But the problem was: Kodak had a big film business - with a lot of film factories, a lot of employees, a lot of executives, and a lot of recurring revenue. And jumping into digital with both feet would have threatened all that. So they didn't capitalise on their early lead - and now they're bankrupt, reduced to licensing their brand to third-party battery makers. > You can display relevant ads in LLM output or natural language web-search. Maybe. But the LLM costs a lot more per response. Making half a cent is very profitable if you only take 0.2s of CPU to do it. Making half a cent with 30 seconds multiple GPUs, consuming 1000W of power... isn't.
- djtango 1y agoThis is a good anecdote and it reminds me of how Sony had cloud architecture/digital distribution, a music label, movie studio, mobile phones, music players, speakers, tvs, laptops, mobile apps... and totally missed out on building Spotify or Netflix. I do think Google is a little different to Kodak however; their scale and influence is on another level. GSuite, Cloud, YouTube and Android are pretty huge diversifications from Search in my mind even if Search is still the money maker...
- danpalmer 1y agoThis would be like claiming in 2010 that because Page Rank is out there, search is a solved problem and there’s no secret sauce, and the following decade proved that false.
- noosphr 1y agoIn a time where statistical models couldn't understand natural language the click stream from users was their secret sauce. Today a consumer grade >8b decoder only model does a better job of predicting if some (long) string of text matches a user query than any bespoke algorithm would. The only reason why encoder only models are better than decoder only models is that you can cache the results against the corpus ahead of time.
- jampekka 1y ago> Modern BERT with the extended context has solved natural language web search. I doubt this. Embedding models are no panacea even with a lot simpler retrieval tasks like RAG.
- noosphr 1y agoRAG is literally what Google Search is. Unlike the natural language queries that RAG has to deal with, Google searches are (usually) atomic ideas and encoder-only models have a much easier time with them.
- marsten 1y agoI think what may save Google from an Innovator's Dilemma extinction is that none of the AI would-be Google killers (OpenAI etc.) have figured out how to achieve any degree of lock-in. We're in a phase right now where everybody gets excited by the latest model and the switching cost is next to zero. This is very different from the dynamics of, say, Intel missing the boat on mobile CPUs. I've been wondering for some time what sustainable advantage will end up looking like in AI. The only obvious thing is that whoever invents an AI that can remember who you are and every conversation it's had with you -- that will be a sticky product.
- noosphr 1y agoWho ever gets AI to be able to search the whole corpus of human knowledge. I'm not just talking about web pages, I'm talking every book, every scientific paper, every news paper, every piece of text stored somewhere. I've build RAG systems that index tokens in the 1e12 range and the main thing stopping us from having a super search that will make google look like the library card catalogue is the copyright system. A country that ignores that and builds the first XXX billion parameter encoder only model will do for knowledge work what the high pressure steam engine did for muscle work.
- krackers 1y agoAssuming that DeepSeek continues to open-source, then we can assume that in the future there won't be any "secret sauce" in model architecture. Only data and training/serving infrastructure, and Google is in a good position with regard to both.
- fulafel 1y agoMaking your own hardware would seem to yield freedoms in model architectures as well since performance is closely related to how the model architecture fits the hardware.
- jononor 1y agoGoogle is also in a great position wrt distribution - to get users at scale, and attach to pre-existing revenue streams. Via Android, Gmail, Docs, Search - they have a lot of reach. YouTube as well, though fit there is maybe less obvious. Combined with the two factors you mention, and the size of their warchest - they are really excellently positioned.
- mattlondon 1y agoYouTube is very well positioned - all these video generating models etc. I am sure they'll be loads of AI editors too
- vitaflo 1y agoGood, maybe Youtube will finally recommend something to me I actually want to watch.
- HaZeust 1y agoPersonally, I've never actually heard this problem. Do you watch industry-specific videos in a non-anonymized browser session enough? Once you watch, like, 5 videos on topics you care about, the algorithm has no shortage of astute suggestions.
- retinaros 1y agothey re not alone to do that tho.. aws also does and I believe microsoft is into it too
- marcusb 1y agoFrom the article: > I’m forgetting something. Oh, of course, Google is also a hardware company. With its left arm, Google is fighting Nvidia in the AI chip market (both to eliminate its former GPU dependence and to eventually sell its chips to other companies). How well are they doing? They just announced the 7th version of their TPU, Ironwood. The specifications are impressive. It’s a chip made for the AI era of inference, just like Nvidia Blackwell
- thunderbird120 1y agoNice to see that they added that, but that section wasn't in the article when I wrote that comment.
- marcusb 1y agoMaybe they read your comment?
- SubiculumCode 1y agoIt was there.
- marcusb 1y agoTo be fair to thunderbird120, the author of this piece made edits at some point. See https://archive.is/K4n9E https://archive.is/K4n9E. No discussion of the recent TPU releases, or TPUs for all, for that matter.
- SubiculumCode 1y agoYou are correct. I misjudged.I thought I had read the article early, it must have been just after the edits.
- marcusb 1y agoYou and me both.
- imtringued 1y agoGoogle is what everyone thinks OpenAI is. Google has their own cloud with their data centers with their own custom designed hardware using their own machine learning software stack running their in-house designed neural networks. The only thing Google is missing is designing a computer memory that is specifically tailored for machine learning. Something like processing in memory.
- ENGNR 1y agoThe one thing they lack that OpenAI has is… product focus. There’s some kind of management issue that makes Google all over the shop, cancelling products for no reason. Whereas Sam Altmans team is right on the money. Google is catching up fast on product though.
- mike_hearn 1y agoTPUs aren't necessarily a pro. They go back 15 years and don't seem to have yielded any kind of durable advantage. Developing them is expensive but their architecture was often over-fit to yesterday's algorithms which is why they've been through so many redesigns. Their competitors have routinely moved much faster using CUDA. Once the space settles down, the balance might tip towards specialized accelerators but NVIDIA has plenty of room to make specialized silicon and cut prices too. Google has still to prove that the TPU investment is worth it.
- dgacmu 1y agoThey go back about 11 years.
- phillypham 1y agoDepending how you count, parent comment is accurate. Hardware doesn't just appear. 4 years of planning and R&D for the first generation chip is probably right.
- dgacmu 1y agoThe first TPU (Seastar) was designed, tested, and deployed in 15 months: https://arxiv.org/pdf/1704.04760 https://arxiv.org/pdf/1704.04760 They started becoming available internally in mid 2015.
- mike_hearn 1y agoI was wrong, ironically because Google's AI overview says it's 15 years if you search. The article it's quoting from appears to be counting the creation of TensorFlow as an "origin".
- dgacmu 1y agoThat's awesome. :) and even that article is off. They probably were thinking of DistBelief, the predecessor to TF.
- 1y ago
- albert_e 1y agoAmazon also invests in own hardware and silicon -- the Inferentia and Trainium chips for example. But I am not sure how AWS and Google Cloud match up in terms of making this verticial integration work for their competitive advantage. Any insight there - would be curious to read up on. I guess Microsoft for that matter also has been investing -- we heard about the latest quantum breakthrough that was reported as creating a fundamenatally new physical state of matter. Not sure if they also have some traction with GPUs and others with more immediate applications.
- chazeon 1y agoI think Amazon, Meta have been trying on inference hardware, they throw their hands up on training; but TPUs can actually be used in training, based on what I saw in Google’s colab.
- jxjnskkzxxhx 1y agoI've used Jax quite a bit and it's so much better than tf/pytorch. Now for the life of me, I still haven't been able to understan what a TPU is. Is it Google's marketing term for a GPU? Or is it something different entirely?
- JLO64 1y agoTPUs (short for Tensor Processing Units) are Google’s custom AI accelerator hardware which are completely separate from GPUs. I remember that introduced them in 2015ish but I imagine that they’re really starting to pay off with Gemini. https://en.wikipedia.org/wiki/Tensor_Processing_Unit https://en.wikipedia.org/wiki/Tensor_Processing_Unit
- jxjnskkzxxhx 1y agoBelieve it or not, I'm also familiar with Wikipedia. It reads that they're optimized for low precisio high thruput. To me this sounds like a GPU with a specific optimization.
- flebron 1y agoPerhaps this chapter can help? https://jax-ml.github.io/scaling-book/tpus/ https://jax-ml.github.io/scaling-book/tpus/ It's a chip (and associated hardware) that can do linear algebra operations really fast. XLA and TPUs were co-designed, so as long as what you are doing is expressible in XLA's HLO language (https://openxla.org/xla/operation_semantics https://openxla.org/xla/operation_semantics), the TPU can run it, and in many cases run it very efficiently. TPUs have different scaling properties than GPUs (think sparser but much larger communication), no graphics hardware inside them (no shader hardware, no raytracing hardware, etc), and a different control flow regime ("single-threaded" with very-wide SIMD primitives, as opposed to massively-multithreaded GPUs).
- jxjnskkzxxhx 1y agoThank you for the answer! You see, up until now I had never appreciated that a GPU does more than matmuls... And that first reference, what a find :-) Edit: And btw, another question that I had had before was what's the difference between a tensor core and a GPU, and based on your answer, my speculative answer to that would be that the tensor core is the part inside the GPU that actually does the matmuls.
- deleted 1y ago[deleted]
- acstorage 1y agoUnclear if they can actually beat GPUs in training throughout with 4D parallelism
- 6510 1y agoThe problem is always their company never the product. They had countless great products. You cant depend on a product if the company is reliably unreliable enough. If they don't simply delete it for being expensive and "unprofitable" they might initially win, eventually, like search and youtube, it will be so watered down you cant taste the wine.
- AlbertoRomGar 1y agoI am the author of the article. It was there since the beginning, just behind the paywall, which I removed due to the amount of interest the topic was receiving.