7 ms·
The Tesla Dojo Chip Is Impressive, but There Are Some Major Technical Issues
- rektide 5y agosome discussion yesterday, https://news.ycombinator.com/item?id=28361807 https://news.ycombinator.com/item?id=28361807 it's interesting because it's clearly exciting & leading edge tech. unlike most Tesla tech which ultimately has consumers using it, where we all get to assess strengths & weaknesses, this tech is going to remain inside the Tesla castle, unviewable, unassessable. we'll probably never know what real strengths or weaknesses it has, never understand all the ways it doesn't well, or as well as competitors. it's going to remain an esoteric dollop of computing.
- cheinic293 5y ago> we'll probably never know what real strengths or weaknesses it has, never understand all the ways it doesn't well, or as well as competitors. Sounds like a classic Elon play: 1. Hype up FSD 2. First time Tesla buyers fork over an additional $10K 3. Keep delaying FSD by a decade 4. Stock keeps going up
- gootler 5y agoAnother butt munch in their momma's basement writing a hit piece because Tesla isn't a union shop. So easy to spot these losers. They think they're so smart yet they can only write shit on a blog rather than invent something useful.
- CasillasQT 5y ago"We believe it makes sense for Tesla to pour as much capital as needed into winning the Robotaxi race and catch up to these two". That has to be a joke right?
- rvz 5y agoAt this point, everyone knows the whole thing is a joke. The robotaxi race was supposed to be already finished by 2020 alongside with FSD at Level 5 - today it is still Level 2. So, where are the so-called robotaxis?
- CasillasQT 5y agoDid you see the presentation from Karpathy? Tesla goes for a general vision only end-to-end deep AI model that could in theory get rolled out everywhere on earth with enough training and a good approach for fast edge-case solving, which they showed how this can be accomplished. All the other players try to solve this with lidar and cars that cost around 500k to build and they have pretty much 0 data except for the maps they generate themselves. This approach will never solve L5. Tesla may need another 10 years, but they are so far out of reach of the other players that you cant even call them competition at this point.
- eddanger 5y agoFolks understand L1 and L5. The levels between are such a blur and mix-match of things that I don't think anything is even accomplished having these levels. I agree with you, from what I have seen Tesla is far ahead and rapidly progressing using the right approach of training NNs while having a human behind the wheel.
- mlindner 5y agoThe levels are even faulty. They assume that geofencing is required and so geofencing is hard coded into the levels.
- yourapostasy 5y ago> ...the right approach of training NNs while having a human behind the wheel. Can someone please help me across a conceptual bridge here? Is there some work I'm not familiar with that shows humans use the biologically-equivalent NNs used by Tesla to accomplish L5-grade driving? I'm not talking about doing it quickly, I'm at this point interested in Tesla or anyone else for that matter demonstrating doing it at all, at any speed. It can be at an agonizingly-slow 0.25 km/hour and that would be fine. I'm having trouble bridging between L5-the-destination and NNs-are-definitely-the-way-to-get-there. This sounds an awful lot like saying NNs-are-the-Moravec's-Paradox-solution, and I'm not sure I've read conclusively how that can be true. I can accept it as a hypothesis, but other than actually trying it out like Tesla is doing, I haven't read why it is such a strong conjecture. It sounds from articles like [1] and [2] Tesla is only just now starting to really get into applying NNs more broadly to the problem space, and the prior years were mostly focusing on more conventional machine vision techniques and getting clean data for NNs to ingest. But I've yet to read a convincing explanation for how ML will functionally solve even the subset of Moravec's Paradox needed to accomplish L5. I grant that it will solve a facsimile of the paradox, but I feel it is arguable if it will be a reasonable facsimile. That sounds an awful lot like, "we'll brute force throw enough training data at it to reach 'reasonable facsimile' level", and I'm cautious when I hear of brute forcing as a strategy for arriving at R&D results. [1] https://insideevs.com/news/466239/tesla-migrating-to-neural-net-self-driving-decisions/ https://insideevs.com/news/466239/tesla-migrating-to-neural-... [2] https://electrek.co/2021/02/08/tesla-looks-hire-data-labelers-feed-autopilot-neural-nets-images-gigafactory-new-york/ https://electrek.co/2021/02/08/tesla-looks-hire-data-labeler...
- m3kw9 5y agoAll I see is [techno terms].. impressive engineering.. lots of problems need to be solved first..2022..on paper toe to toe with Nvidia.. calm the hype.
- neolefty 5y ago> This chip is not Tesla designing something that is better than everyone else all by themselves. We are not at the liberty to reveal the name of their partner(s), but the astute readers will know exactly who we are talking about when we reference the external SerDes and photonics IP. Any "astute readers" here who know who the partner would be?
- eddanger 5y agoMy astute guess is ASML.
- TomVDB 5y agoThat astute guess is so out of left field that it needs at least a bit of justification. To my limited ASML knowledge, ASML is a fab equipment maker, not a silicon IP procider.
- gimmeThaBeet 5y agoCorrect, that's a bit of an odd guess. You are definitely working closely with your foundry on something like this, but unless you are building some truly exotic device, I don't think their vendor is getting involved. Like Melkman says, Broadcom is a good guess. Not only for past rumors, but iirc they also did work with google's TPU (could never figure out if that was actually confirmed?). Interconnect IP like that is definitely in their wheelhouse.
- Melkman 5y agoI'd guess Broadcom.See also https://www.broadcom.com/info/optics/silicon-photonics https://www.broadcom.com/info/optics/silicon-photonics
- wumpus 5y agoLooks like the tech that Intel has been attempting to get to work and be cost effective for several decades.
- ggoo 5y agoTesla's claim to delivery ratio is abysmal. I'm not sure why anybody even bothers deconstructing these presentations anymore, they're just fluff.
- snorrah 5y agoI would argue it’s always useful to see their tech deconstructed and explained. If nothing else, so we get an idea what the reality is to counter possible outlandish claims from overly-enthusiastic followers of the company (and its CEO)
- stcredzero 5y agoTesla's claim to delivery ratio is abysmal. Can you substantiate this concretely? How about a list, with direct sources? (Not opinion pieces.)
- ggoo 5y agohttps://en.wikipedia.org/wiki/Criticism_of_Tesla,_Inc https://en.wikipedia.org/wiki/Criticism_of_Tesla,_Inc.
- gooseus 5y ago> The full system is scheduled for some time in 2022. Knowing Tesla’s timing on Model 3, Model Y, Cyber Truck, Semi, Roadster, and Full Self Driving, we should automatically assume we can pad this timing here. That's just from the article; off the top of my head: * NYC to LA fully autonomous drive by 2017. * 1M Robotaxis on the road by 2021. * Hyperloop. * Solar roof tiles. * All superchargers will be solar-powered. * Tesla Semi. Sure, some of these things may be "just around the corner" or "ramping up now", but some of these are claims going back almost 5 - 10 years where Elon says "2 weeks", "next year", "2 years", really whatever it takes to be just believable enough to get enough people to buy into a future where Tesla is worth 10x what it is today.
- stcredzero 5y agoSo, to sum up, the criticisms are mostly about "Elon Time" and not about whether Tesla is actually trying to do things. Hyperloop is not something Tesla is trying to do, and Elon is only cheerleading that effort. There's something highly off-kilter with the relative mildness of the above and the vitriol of the criticism directed at it. This actually makes me feel really good about Elon Musk and Tesla's prospects!
- dragontamer 5y agoSomehow, I'm reminded of the Tsar tank from WW1. The Russians knew that a new weapon of war: an armored car, was necessary to break the stalemate of trench warfare. This hypothetical armored car needed many features: the most important was that it must be able to move across the muddy no man's land reliably. Tests have shown that regular sized wheels would get stuck in the mud. A bigger wheel has more surface area and greater contact area. So the Russians built an armored car with the largest wheels possible. Russian tests were outstanding, the Tsar tank rolled over a tree !!!! https://en.m.wikipedia.org/wiki/Tsar_Tank https://en.m.wikipedia.org/wiki/Tsar_Tank The French design was to use caterpillar tracks. We know what works now since we have a century of hindsight. -------- Spending the most money to make the biggest wheel isn't necessarily the path to victory. I think it's more likely that the tech (aka, caterpillar track equivalent) hasn't been invented yet for robotaxis. Hitting the problem with bigger and more expensive neural network computers doesn't seem to be the right way to solve the problem.
- deleted 5y ago[deleted]
- zaptrem 5y agoI agree with your points on the robotaxi front, but there are many other problems that will totally benefit from a bigger training computer.
- baybal2 5y ago> many other problems that will totally benefit from a bigger training computer. I don't really think it's that many. The industry collectively sank untold billions into the blind belief that neural algorithms will somehow turn into "AI." 10 years later, no "AI," and not even a single money making niche use. Right now the industry is deep in sank cost falacy, and people who promised this, and that to investors are now desperate, and doubling their bets in hopes that "at least something will come out of it...," a casino mode basically.
- gpm 5y ago
- justapassenger 5y ago> Of this competition, only Google and Nvidia have supercomputers that stand toe to toe with the Tesla’s Even assuming that it's true (which I very much doubt - anyone that's willing to spend enough money with Nvidia, can have powerful supercomputer fairly quickly), it's very dishonest statement. It's comparing deployed system with a lab prototype of a single competent of potential supercomputer, that may be fully operational in few years (software is a really, really, really big deal here).
- jeffbee 5y agoIt is really unreasonable to compare Tesla's photoshop mocks with hardware already deployed in the field today. Google already has a TPUv4 cluster that can train ResNet-50 in 13 seconds, which is ridiculous. Until Tesla publishes actual MLPerf benchmarks, you can assume that their ASIC game is at least as far behind Google's as their self-driving game is behind Waymo's: 5 years at a minimum. https://github.com/mlcommons/training_results_v1.0/tree/master/Google/systems https://github.com/mlcommons/training_results_v1.0/tree/mast...
- solidasparagus 5y agoI assumed they're talking about Tesla's A100 cluster, which is huge - https://blogs.nvidia.com/blog/2021/06/22/tesla-av-training-supercomputer-nvidia-a100-gpus/ https://blogs.nvidia.com/blog/2021/06/22/tesla-av-training-s... Tesla's compute-to-researcher ratio is definitely rare
- lpapez 5y ago2022 will surely be the year of Linux on Desktop and fully self driving cars.
- michelpp 5y agoClearly a a shot across the bow for Cerebras and another excellent target for the GraphBLAS. Dense numeric processing for image recognition is a key foundation for what Tesla is trying to do, but that tagging of the object is just the beginning of the process, what is the object going to do? What are its trajectories, what is the degree of belief that a unleashed dog vs a stationary baby carriage is going to jump out? We are just beginning to scratch the surface of counterfactual and other belief propagation models which are hypersparse graph problems at their core. This kind of chip, and what Cerebras are working on, are the future platforms for the possibility of true machine reasoning.
- jvanderbot 5y agoAs Hamming suggested in "Art of doing science and engineering", when you want to make something autonomous, you usually have to build a completely different device that solves the same problem, rather than automating the same device. I wonder. For all the money thrown into self-driving cars research, could we have had an autonomous rail system by now? The technology for mostly-autonomous rail is well understood. Most of the financial cost is in infrastructure to support the system. Seems to me self-driving cars try to short-circuit that infrastructure build-up. They try to "automate the device" rather than "producing an automated system that solves the problem of moving people and goods". Specifically, I wonder if, for the cost and time spent on CPU-and-engineer-driven research and development of autonomous cars, if we could have had nationwide autonomous rail rolled out by now.
- dragontamer 5y ago> could we have had an autonomous rail system by now? We already have autonomous rail systems. Its called positive train control and was fully implemented like a year or two ago (mandated in 2009, but you know how government works, lol) https://en.wikipedia.org/wiki/Positive_train_control https://en.wikipedia.org/wiki/Positive_train_control The train conductor has become more-and-more automated to remove the chance of human error. It works with a system of very reliable sensors that indicate where every train engine is on the rails. Given the huge amount of cargo any particular train has, I don't think there's any intent on cutting the last two humans (the conductor + engineer) out of their job. Their salary costs are miniscule compared to the safety value they deliver, even if the job of driving a train has been almost entirely automated away by now.
- jvanderbot 5y agoI wonder if, for the cost spent on CPU-and-engineer-driven research and development of autonomous cars, if we could have had nationwide autonomous rail rolled out.
- dragontamer 5y agohttps://railroads.dot.gov/train-control/ptc/positive-train-control-ptc https://railroads.dot.gov/train-control/ptc/positive-train-c... > On December 29, 2020, FRA announced that PTC technology is in operation on all 57,536 required freight and passenger railroad route miles, prior to the December 31, 2020 statutory deadline set forth by Congress. We got that. We literally got that.
- danso 5y agoI confess I have a reflexive skepticism to the idea that Tesla's achievements (and struggles) in car manufacturing would translate to any kind of lead in chip design and manufacturing. How long did it take Apple from planning to rollout for M1? And the Tesla chip seems to be making bigger revolution-sized claims?
- boardwaalk 5y agoTesla is already shipping their own chips in every car. And that’s a better comparison (it’s an end user thing you can buy) than this data center processor. It’s hard to compare vs, say, Nvidia’s car computers because it’s all locked down. But I believe the energy efficiency is fairly good.
- scardycat 5y agoThis is a step in the right direction. I witnessed the semiconductor industry abandoning their own designs in favor of Intel/x86. Better diversity in chip design is always a good thing, even if its in closed ecosystems (Google TPU, Tesla Dojo)
- modeless 5y agoI'm glad people are exploring the design space. To some extent the training techniques and neural net architectures need to be tailored to the hardware. Nvidia isn't on top just because they're good at chip design, but because people have chosen to focus research effort on techniques that work well on Nvidia hardware. New hardware may allow new techniques to shine. New hardware architectures can't really be used to their full potential without years of research into techniques that are suited for them. The more people who have access to the hardware, the faster we can discover those techniques. If Tesla is serious about their hardware project, they need to offer it to the public as some kind of cloud training system. They don't have enough people internally to develop everything themselves in a short enough time to remain competitive with the rest of the industry.
- 2bitencryption 5y agofrom the article: > but the short of it is that their unique system on wafer packaging and chip design choices potentially allow an order magnitude advantage over competing AI hardware in training of massive multi-trillion parameter networks. I kind of wonder if Tesla is building the Juicero of self-driving. [0] Beautifully designed. An absolute marvel of engineering. The result of brilliant people with tons of money using every ounce of their knowledge to create something wonderful. Except... you could just squeeze the bag. You could just use LIDAR. You could just use your hands to squish the fruit and get something just as good. You could just (etc etc). No doubt future Teslas will be supercomputers on wheels. But what if all those trillions of parameters spent trying to compose 3D worlds out of 2D images is pointless if you can just get a scanner that operates in 3D space to begin with?? [0] https://www.theguardian.com/technology/2017/sep/01/juicero-silicon-valley-shutting-down https://www.theguardian.com/technology/2017/sep/01/juicero-s...
- arnaudsm 5y agoThe Juicero comparison doesn't hold up. LIDAR is 10x more expensive than RGB, but neither reach lvl5 at the moment. I'm glad multiple companies try multiple paths, it's the best way to avoid a research dead-end.
- KaiserPro 5y ago> LIDAR is 10x more expensive than RGB but pure RGB needs $millions to make a reliable realtime depth sensor, plus custom silicon and a massive annotated dataset. It might just be that one company can do it, but its a hefty gamble.
- nightski 5y agoEveryone acts like LIDAR is the holy grail but then why isn't there someone destroying Tesla with that tech? Waymo is not much farther along than Tesla, maybe even behind as far as miles driven. If that was all that was needed then it would be done.
- thesausageking 5y agoThe Q&A section on their compiler and software that the author links to is very interesting: https://www.youtube.com/watch?v=j0z4FweCy4M&t=8047s https://www.youtube.com/watch?v=j0z4FweCy4M&t=8047s It sounds like they're going to have write a ton of custom software in order to use this hardware at scale. And, based on the team being speechless when asked a follow up question, it doesn't sound like they know (yet) how they're going to solve this. Nvidia gets a lot of credit for their hardware advances, but what really what their chips work so well for deep learning was the huge software stack they created around CUDA. Underestimating the software investment required has plagued a lot of AI chip startups. It doesn't sound like Tesla is immune to this.
- cr4zy 5y agoTrillion parameter networks are mentioned a few times, but Tesla is deploying much smaller networks than that (like tens of millions IMU). Trillion param networks are mostly transformers like GPT-3 (actually 175B) etc... that are particularly heavy vs Conv as they have no weight sharing. Tesla is definitely starting to use transformers though, e.g. for camera fusion and evidenced by their focus on matrix multiply in dojo asic's vs the conv asics they have in the on-vehicle chips.
- zozbot234 5y agoYup, there's plenty of ML architectures that try to save on parameters size, achieving better generalization (less overfitting) at the expense of slightly costlier training and inference. The memory constraints on Tesla Dojo might not be a big deal after all.
- immmmmm 5y agonot fully related but i was doing some reading on various "new sustainable ways of transportation" and, since they're building the biggest hyperloop test track near my place, i found this interesting video of some of problems one might get trying to put vacuum in a pipe: https://youtu.be/Zz95_VvTxZM https://youtu.be/Zz95_VvTxZM
- Const-me 5y ago> they have 1.25MB of SRAM and 1TFlop of FP16/CFP8… This is woefully unequipped for the level of performance they want to achieve. Any idea how OP made that conclusion? My GeForce 1080Ti has 1.3MB of in-core L1 caches (28 streaming multiprocessors, 48kb L1 each). It also has L2 but not too large, slightly under 3MB for the whole chip. The GPU delivers about 10 TFlops of FP32 which needs 2x the RAM bandwidth of FP16. I’m generally OK with the level of performance, at least until the GPU shortage is fixed.
- thunkshift1 5y agoWhat a bs fanboy article.. the author is going gaga over something that isnt even out in silicon yet, and has no credible plans of software ecosystem coming on top the hw( if it materializes). Unbelievable hype.
- _nalply 5y agoMy curiosity got piqued at the mention of CFP8 (configurable floating point 8), but googling this didn't yield usable information. What exactly is CFP8? How many bits does one instance of CFP8 use? What mathematical operations are supported? How does one configure the floating point?
- _nalply 5y agoI found about posits. https://www.johndcook.com/blog/2018/04/11/anatomy-of-a-posit-number/ https://www.johndcook.com/blog/2018/04/11/anatomy-of-a-posit... Perhaps CFP8 are parameterized 8-bit posits where the parameter is the value es. The larger es is, the greater the dynamic range is at the expense of precision. Two examples: posit<8, 0> (es = 0) has as largest positive number 64 and the smallest positive number 1/64. posit<8, 1> (es = 1) has as largest positive number 4012 and the smallest positive number 1/4012. The formula for the largest positive number for 8-bit posits is: 2 ^ 2 ^ es ^ 6. posits don't have NaNs and only one infinity (±∞), so they can use more of the 8 bit values for numbers than floating point numbers. I wonder: is CFP8 = posit<8, es>?