9 ms·
OpenAI Jalapeño: Better than Nvidia Blackwell
https://www.bloomberg.com/news/articles/2026-08-25/openai-claims-its-new-chips-can-outperform-nvidia-processors-in-tests https://www.bloomberg.com/news/articles/2026-08-25/openai-cl..., https://archive.ph/yCTrr https://archive.ph/yCTrr
- danielovichdk 22d agoI guess special hardware is the new moat in AI. Maybe the money will still flow into this industry after all
- Alien1Being 22d agoWARNING: AI VENDOR HYPE
- ChoosesBarbecue 22d agoThis is most impressive. The interesting question to me, is outside of the LLM accelerator space: will generalized chips have massive leaps in performance once LLM technology is used to create the next generation? In general, will we see rapid advances while we extract the value of these models in creating architectures? I'm so far removed from the space that this is a very naive interpretation of all this, but I'm curious.
- wmf 22d agoExisting CPUs have been extremely optimized by ~6 competing, well-funded teams. I expect AI to accelerate things somewhat but it's not clear that there is any low-hanging fruit available for AI to find.
- manquer 22d agoASICs always do better than general purpose chips. General purpose chips is turtles and turtles of virtualization and have to consider 4+ decades of backward compatible instructions set support. ASICs are deployed when the application area is economically large enough to so there is return on the investment in developing one. Bitcoin mining few years ago or today inference or more mundane things like video decoding/encoding. General purpose chips on the other hand have to be general purpose first to be useful, i.e. support as many application domains and instruction sets as possible . It can be long tail of support which both slow your chip down and also slow development down. Apple's took a long time to develop M series to be general purpose enough and still need even now software tooling like Rosetta to make say virtualization work for a good reason. New tooling would always help and there is already lot of software emulation for developing chips today but you still need physical iterations to tap-out and have high enough yield, no LLM can help with that.
- anthonypasq 22d agoContinued hardware improvements really make it hard for me to believe token prices will not continue to plummet.
- datakan 22d agoToken prices coming down means nothing if the models keep wasting them
- gwerbin 22d agoHopefully this also means billionaires can stop trying to drop data centers into residential neighborhoods with zero noise control and polluting on-site generators, signing local politicians on with NDAs, calling for eminent domain to seize homes to build power lines to data centers, etc. etc. etc. Not to mention the water use controversy. Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.
- tmp10423288442 22d agoNah, Jevon’s Paradox says that cheaper tokens will mean increased overall energy consumption. If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.
- gwerbin 20d agoWe already had plenty of data centers in the US before the AI boom that weren't severely harmful to their neighbors. Cutting red tape is not the same as eliminating meaningful regulation. There are plenty of old industrial sites that could be repurposed as data centers. It turns out it's cheaper to bribe some small town government to give you a tax cut and discounted electricity and water rate.
- vlyan 22d ago>polluting on-site generators how much pollution do you believe modern gas-turbine engines to produce? >Not to mention the water use controversy. what percentage of US water usage do you believe is by AI data centers?
- varispeed 22d agoWhy they don't research how to make their own RAM and they have to buy it from the common market? They should GTFO with this crap. Create barriers to computing for ordinary people while milking businesses for tokens.
- petcat 22d agoBuilding a custom-designed ASIC is much easier than producing state of the art memory chips. There's a reason why Micron and Nvidia are the crown jewels of American technology right now and for the foreseeable future.
- brcmthrowaway 22d agoNVIDIA produces memory?
- fc417fc802 22d agoFabless AFAIK. And that's the actual problem - drawing up CAD diagrams doesn't help if the factories are fully booked out.
- Cyph0n 22d agoA state of the art GPU is much harder to design & produce at scale and than an internal ASIC.
- JV00 22d agoNvidia does not make RAM
- varispeed 22d agoThat doesn't excuse them from wrecking the market for ordinary person.
- chris_money202 22d agoNvidia buys the memory it uses on its GPUs, same as all other ASICs. To give some context, Intel started making DRAM, I think they were actually the company that came up with modern memory techniques. They exited the market and pursued a more lucrative moat with CPUs.
- epistasis 22d agoIt's so funny to see FP4.... I remember 20 years ago being asked what sort of HPC we needed in genomics, and the answer was basically, "lower precision, faster" for the stuff I was working on. But FP4 is, well, almost comical. One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)
- nxtfari 22d agoAgree, I remember when even half precision made its way into C# sometime around 2020 (I didn’t know much about ML then) and I thought, well I guess that’s a worthwhile tradeoff but I can’t imagine going lower. Lo and behold (1-bit Bonsai) how much lower you could go.
- jimmySixDOF 22d agoI love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al
- xyzsparetimexyz 22d agos** posting? sex posting?
- msh 22d agoshit posting
- minimaltom 22d agoThats what I thought too but then it would be s**?
- jareklupinski 22d agoi see 'hunter2'
- madspindel 22d agos**? Edit: OK, hn is removing one *
- yjftsjthsd-h 22d agoIf it's trying to convert it to italics, you may have to use a backslash to escape them
- masfuerte 22d agoOr double them up: s****** gives s***.
- empath75 22d agoWhen people talk about the commodification of inferencing, they imagine a future where everyone has access to frontier models and can run them at the same cost, and what will actually happen is closer to the commodification of _oil_, where only a few companies have the scale to produce it at a competitive price, and advances like this are _why_. Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.
- simianwords 22d agoI don't believe models will be commodified because each model is unique with strengths and weaknesses. Its not like Steel which is more or less the same no matter where you purchase it from. If what you said were true, you would hardly see people complaining about the quality of Opus 5 or good writing from Sol. But people do.
- airspresso 22d agoThis depends heavily on what the use-case is. Yes, if it's a coder making software and having to read LLM output then writing style matters. If the LLM is used in an automated data processing pipeline with a capped level of complexity, entirely different aspects matter and LLMs become more interchangeable.
- lelanthran 22d ago> I don't believe models will be commodified because each model is unique with strengths and weaknesses. They are all converging.
- zurfer 22d agoThe same level of intelligence gets roughly 10x cheaper per year. So you might both be correct where a large part are commodity tasks but frontier is hard and valuable and not commodities.
- skhameneh 22d ago> Its not like Steel which is more or less the same no matter where you purchase it from. I’m not an expert in metallurgy by any means, but this seems really off. There are many recipes for steel and varied processes that also impact the final product.
- fraboniface 22d agoI hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.
- danishanish 22d agoI mean, surely when quality is accounted for the difference is significantly higher
- GaggiX 22d agoOr maybe significantly lower.
- plasticchris 22d agoProbably not when you consider the training cost and upkeep expenses, not to mention the depreciation…
- Phemist 22d agoThe 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.
- phoghed 22d agokind of a moot point if you can't get your brain to not do everything else. I think it's a fun comparison, even if it's not a 100% equivalence.
- CooCooCaCha 22d agoAnd the brain is literally only producing electrochemical signals. I don’t see how tokens can’t produce speech or track metabolic needs. You can talk to chatgpt can’t you? Or do you mean literally talking? Because that’s not a brain function, that’s the mouth, vocal chords, and lungs.
- LarsDu88 22d agoWell Sam Altman finally has built a moat against Chinese open weight AI. Well done. But what will this mean for Cerebras? I remember when Tesla was building its own inference chips, and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies. I suspect the same will be the case with OpenAI vs Cerebras + Nvidia/Groq
- brcmthrowaway 22d ago[dead]
- Eridrus 22d agoCerebras is targeting a distinctly different point on the cost/latency curve. They are betting that there will be some high value applications where latency and not just throughput is super important.
- porridgeraisin 22d agoIt is being used as part of a combined system. For example AWS is pushing for Trainium + WSE 3. The WSE 3 does the decode and the Trainium does the prefill. Even in nvidia land rubin + LPU does a similar thing. It has its downsides of course - if your traffic swings prefill heavy to decode heavy, you can't suddenly use your lpu for prefill. With GPUs they're totally interchangeable. Tradeoffs.
- Eridrus 22d agoAFAIK You can use WSE/LPU for prefill, it's just less efficient to do so.
- porridgeraisin 22d agoWell ya, that efficiency is why it's split. There is also the other idea where you run your attention layer on the GPU/TPU/Trainium and the FFN on the SRAM accelerator. Because KV cache is more difficult on cerebras etc, while MOE latency is easier to deal with
- simianwords 22d agoHow can OpenAI mass produce this chip at scale more economically than Nvidia which has experience in the supply chain and scale efficiencies to do it efficiently?
- airspresso 22d agoBy leveraging the experience Broadcom has in this area. Still remains to be seen how that goes when they want to scale production.
- chris_money202 22d agoIn the short and medium term, it probably won't be more economical to produce for OpenAI. Where OpenAI is benefitting from their own chip is being able to tailor it to their models and workloads. When you buy off the shelf Nvidia, its not perfectly tailored and OpenAI has to spend marginally more to run off that chip. At the scale OpenAI is operating at and plans to operate at, that margin becomes pretty big $$
- dpe82 22d agoNVidia has enormous operating margins, so a competitive solution doesn't have to match or beat NVidia's scale efficiencies; it just has to beat delivered cost. One objective of the project might be simply to provide credible negotiating leverage when dealing with existing suppliers like NVidia. You don't have to deploy at scale for that to work, but you do have to look like you could if pushed hard enough.
- vntok 22d ago> NVidia has enormous operating margins, so a competitive solution doesn't have to match or beat NVidia's scale efficiencies; it just has to beat delivered cost. But then that means you have no actual moat against the behemot, right? Your competitor can move into the market as soon as they want to, at much better cost (so at slightly better price)... and Nvidia certainly can adapt much faster around hard hardware specs innovation than a new entrant ever could.
- 22d ago
- thebeardisred 22d agoAll of these words spilled and no mention of the ISA.
- 0xbadcafebee 22d agoStory says they're power limited. That's half-true. Actually they're water-limited. To generate power, you need water. To cool chips, you need water. If you try to use less water on one side, you need more water on the other side (it's physics ya'll, making and using energy generates heat which requires dissipation). The world's freshwater is diminishing while also being consumed at an alarming rate. The future AI oligarchs are whoever controls the most water. The other side of the conversation is the idea that large models in DCs on custom silicon is the future. Maybe for enterprise? But consumers will eventually (10 yrs) have affordable hardware designed to run crazy-good local models (more RAM + higher bandwidth). That will take pressure off of datacenters, but also reduce AI profits, and move that money to consumer chip/device makers. Apple is once again the biggest winner. Nvidia consumer chips might get cheaper, but nerfed, to encourage datacenter use where they make more money. I'm hoping AMD can stop being terrible at software so that when we finally have their better hardware we can actually use it.
- minimaltom 22d agoFor datacenters specifically I've never understood what specifically consumes the water. Arent the water-cooling loops closed, so the water just cycles around and around and around?
- justincormack 22d agoYes they are for water cooling.
- SirMaster 22d agoThey evaporate the water which is what makes it cool so effeciently.
- Ductapemaster 22d agoEvaporative cooling does not necessitate an open loop system
- 22d ago
- throwaw12 22d agoCompetition is good for all of us, we will get better and faster chips. Or at least Nvidia GPUs will become slightly cheaper for regular consumers again
- theandrewbailey 22d agoThe pricing of GPUs themselves aren't really the problem: it's the VRAM that comes with them.
- WarmWash 22d agoThat's if any datacenters are allowed to be built with them. There is probably a ~50% chance that the next Dem candidate for presidency runs on a national datacenter moratorium or something equally as crippling.
- bigyabai 22d agoIf the populist campaign is to Make Affordable DRAM Again, then it's not a terrible solution. The current datacenter owners love a compute-bound world anyhow. A moratorium on new datacenters would increase their valuation, encourage efficiency and make computers cheap again. If Chinese labs can ship frontier models under 1T parameters, why not American labs too?
- porridgeraisin 22d agoThese are not replacing GPUs, they are entirely complementary. It's the same with cerebras, groq etc, they are all complementary to the GPU.
- einpoklum 22d agoIf you think tanking Trillions in investments, warming the earth and increasion ocean water levels, creating water shortages and brown-outs is "good for all of us" - well, the rest of us beg to differ.
- throwaw12 22d agothese GPUs make computation faster, I understand as of now maybe all the computation is used to generate yet another junk LinkedIn post or unnecessary RFC, but at some point this craze should settle and we will be left with powerful computation machines, which can be used for computing more useful things
- lelanthran 22d agoThis means that they're going to want to IPO soon - this is good news for investors + they need the capital.
- rsync 22d agoNo, this is because they want to IPO soon. If the chips weren't this compelling they would have something different to announce. These are paperclip maximizers who just happen to wear human skin - there is no underlying premise nor ideological goal.
- mchusma 22d agoI think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves. For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough. While 2 years ago nothing was useful more than 1 year long, there are many older models in use now (e.g. Haiku 4.5, GPT-OSS 120b), and I expect this trend to continue. I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.
- sebzim4500 22d agoMy guess is we only see this once they start saturating computer use benchmarks. That's a use case which would be extremely valuable at the right costs/speed, but the current models just aren't there yet.
- bmulholland 22d agoProbably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this. Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a month, start to finish, for the physical processing. Maybe once LLM improvements asymptote further?
- kushie 22d agotapeout could shrink but days per mask layer (DPML) does not have much margin..
- smallmancontrov 22d agoI'm not in industry, is DPML (which I assume is the time required to make a mask?) set by electron beam scan time or something?
- Alien1Being 22d agoWARNING AI HYPE
- corford 22d agoThese nascent inference chip efforts are reminding me of the early 3dfx / riva / mach / powervr days. Will be interesting to see if inference chips are here to stay and, if so, who the eventual dominant player(s) will be
- ehnto 22d agoWhich in turn reminds me of Soundblaster audio cards! I suspect inference chips are closer to the GPU story than the Soundblaster story though. I remember one soundblaster card I bought came with a Lara Croft demo, that exploited the incredible immersion of real time dynamic reverb. Genuinely I think game audio took a few steps back from that heady era, the innovation in audio likely didn't sell as many cards as graphics innovations did.
- rubzah 21d agoI recently heard, for the first time, what Space Quest sounded like with a Roland board attached. It was mindblowingly amazing. And all most anyone ever experienced was bleep, bloop.
- noir_lord 22d agoOn board got "good enough" and the separate cards died away. In fairness on board (depending on the board but on the whole) is pretty good.
- bayindirh 22d agoEAX was very powerful in its heyday, but it has died because of a thousand cuts. First we had to have the audio processor. Good EAX was available on top of the line cards, and they were not always cheap. Lower end chips got less features. Then we had to have the speaker setup to have the greatest sound, or needed to get a real 5.1 headphones, which were bulky and never provided the same fidelity. Then Microsoft changed the Windows driver model, cutting the driver's direct access to the card. All of the timing sensitive effects were gone in an instant. I remember installing the new drivers and getting literally nothing. Sound Blaster was the only card with an hardware mixer, and Microsoft didn't feel like enabling them. Mixing at the DirectX layer killed the cards. Soundblaster's very closed stance didn't help them either. None of the cards after Audigy2 worked with Linux when I had my desktop system. After my Audigy2ZS, I moved to Asus Xonar D2X. Its positional audio capabilities were nice, but I mostly bought it for its Linux support and sound quality, and that was top notch in that regards. Then sound cards became commodity. Everybody stopped making good cards. Musicians moved to audio interfaces, audiophiles moved to DACs. Just looked to the SoundBlaster website. Internal cards are very limited. One DAC, one DTS enabled 7.1 sound card for PC cinema systems, three game oriented lower end cards, nothing else.
- einpoklum 22d agoI hope the LLM wave will leave GPUs behind to go back to pursue more general-purpose computation rather than spending their die area on multiplying 4-bit-number matrices and such things.
- acedTrex 22d agoIs that not literally the exect opposite of the direction asics for LLM inference is going?
- einpoklum 22d agoMy point is, that if companies develop ASICs for LLM work, then GPUs will stop being the go-to computation device for these workloads, and that will mean, hopefully, that their architectures will stop being warped so as to cater to LLM work.
- mkw5053 22d agoWarning, this is a long comment! (I’m trying to stick to sourced facts here and not overstate what they mean) I went down a rabbit hole after watching Dylan Patel on Dwarkesh today: https://www.youtube.com/watch?v=aV26V1UvkJw https://www.youtube.com/watch?v=aV26V1UvkJw I was initially just surprised by how bullish Dylan is on OpenAI/Anthropic and how bearish he is on China, despite Chinese labs getting closer to US SOTA while offering inference at dramatically lower prices. So, I started digging while waiting for various day-job inference calls to return, ha. Dylan says he spent years obsessively posting on hardware forums, moderating hardware subreddits, and running anonymous hardware blogs/videos before SemiAnalysis. But he also says most of that history is now gone, including from the Internet Archive, because he asked for it to be removed.[1] In a 2024 interview he described his post-college job as “data science” around hurricane/earthquake/wildfire simulations for a financial company.[1] In a 2026 Sequoia interview he described himself as having been a “quant at a small quant risk firm” who generated $10M+ of “risk-free revenue.”[2] The Information reports that he declined to identify the employer and doesn’t list it on LinkedIn.[3] Even harmless/silly stuff seems to drift. In February he said he kept bees for ~1.5 years. Today it was “few months, few months.”[4][5] I know, sort of silly and doesn't matter. The Information reports that Patel owns stakes in ~20 startups in the same ecosystem SemiAnalysis covers, organized a $50M Fluidstack SPV, and is now targeting a $400M venture fund.[3][6] And, in a 2022 HN discussion about SemiAnalysis disclosures, after saying his reports had moved smaller stocks by 20% in a day, Patel wrote: “If I thought I could move the stock, I'd make the position in the morning alongside my clients, and publish shortly after.”[7] I don’t know that any of this is false or that anything improper happened (I’m definitely not claiming that). More that 1-2 of these things would just be odd. Taken together, though, they made me question how much trust I was putting in the broader story. The dynamic of reminds me of crypto, WeWork, Theranos, Citron, etc. Once enough important people validate someone, things that would normally invite basic diligence somehow stop getting questioned. [1] https://www.dwarkesh.com/p/dylan-jon https://www.dwarkesh.com/p/dylan-jon [2] https://sequoiacap.com/podcast/dylan-patel-of-semianalysis-why-hardware-software-co-design-is-ais-real-100x https://sequoiacap.com/podcast/dylan-patel-of-semianalysis-w... [3] https://www.theinformation.com/articles/dylan-patel-semianalysis-grabbed-sway-silicon-valley https://www.theinformation.com/articles/dylan-patel-semianal... [4] https://www.latent.space/p/dylanpatel-cooking https://www.latent.space/p/dylanpatel-cooking [5] https://www.dwarkesh.com/p/dylan-patel-3 https://www.dwarkesh.com/p/dylan-patel-3 [6] https://www.theinformation.com/briefings/exclusive-semianalysis-dylan-patel-targets-400-million-venture-capital-fund https://www.theinformation.com/briefings/exclusive-semianaly... [7] https://news.ycombinator.com/item?id=31065646 https://news.ycombinator.com/item?id=31065646
- a2ff6eeb0 22d agoSounds like a great way to get deals out of Nvidia.
- chabons 22d agoSure, but even heavily discounted Nvidia chips won’t be competitive for inference if they’re worse on perf/W.
- luciana1u 22d ago[flagged]
- calldacopsidgaf 22d agoAny article that features Sam's fucking creepy face should be marked with a jumpscare warning
- tecoholic 22d agoThe reliance on Deepseek and Kimi as the benchmarks from every chip maker from NVIDIA to OpenAI is a good tell of where things are heading. In the next couple of years, hopefully we will have systems at home for everyday use and corporations can buy bulk from providers.
- chabons 22d agoAs opposed to closed-source models? Benchmarks for GPT Sol wouldn’t be particularly meaningful, as no one else can run the benchmark, and we don’t know what the exact model specs are. Picking the best open source models is really the best they can do.
- tecoholic 21d agoFair point. I might be misreading based on my hopes :/
- m4rtink 22d agoSo this will make GPUs and associated affordable for people, rigjt ?
- jpollock 22d agoNo, it's the wafer starts that are driving prices. Switching from Nvidia to custom doesn't change the constraints.
- bjourne 22d agoThe article is a bit naive: > However, as previously mentioned, Jalapeño’s results are obtained without speculative decoding and Vera Rubin’s results use speculative decoding. Speculative decoding leads to a ~3-5x reduction in cost per token. When speculative decoding is implemented on Jalapeño, this will enable Jalapeño to serve tokens even more cost effectively. How much speculative decoding improves throughput is workload-dependent. Yes, it can improve performance by 5x, but it can also slow down performance by 2x. It also requires a vastly more complex decode loop than the standard token-by-token decode. The tech journos didn't ask themselves if speculative decoding improves perf so much why wasn't it on by default? Perhaps, because speculative decoding is not optimal for OpenAI's ASIC? Perhaps that is also why they were only able to benchmark the not-very-representative single-user-decode case?
- dist-epoch 22d agoIn the slides on twitter you can see Jalapeno CAN do speculative decoding. In fact they explicitly mention how compute is disaggregated 3 ways now: prefill, predict, decode, and how a huge Jalapeno advantage is that it uses dark sillicon to switch between these without having to move the KV cache which remains local.
- bjourne 22d agoThen the article contradicts the slides because it states that OpenAI choose not to disaggregate prefill and decode. Idk you men with "predict"---conventional LLM serving comprises only two phases.
- dist-epoch 22d agosorry, my mistake, I meant draft not predict > it states that OpenAI choose not to disaggregate prefill and decode They disaggregate INSIDE the chip, not by having separate machines for the 3 phases. the slides: https://x.com/beffjezos/status/2092416851737518190 https://x.com/beffjezos/status/2092416851737518190
- arrty88 22d agoIs this bad news for Cerebras?
- villgax 21d agoobviously, Anthropic is gonna do their own chips as well just like openai, all hyperscalers already have their own chip stacks & neoclouds wouldnt even carry this as just as they are not carrying AMD stack lol
- 7e 22d agoOnce again we see the classic PR hype machine tactic of comparing a newer chip which is only available as an engineering sample to other chip designs which are widely available and much older. They also fawn over the chip’s TDP when all other chips have to support 16 bit floating point and thus must run much hotter. They make the classic mistake of equating max TDP with in-use-watts, and praise this magnificent (fictitious) performance per watt at FP8 with other chips’ max-TDP at FP16, which draw twice the power. Evidence that the IPO can’t be far away.
- dkhid 22d agoR@6.....111
- dkhid 22d agohacker
- g00afthrowaway 22d agoFunny semi analysis has credibility here of all places. The founder is well-known in the hardware circle to be a black-market information trader. It works like this: 1. Founder befriends undergrad interns/graduate student interns, buys them gifts, invite them to dinner/yacht/house/vc parties etc, or pays them to write articles 2. Founder extracts insider information out of these interns 3. Founder sells this information to companies paying "consulting" fees
- duchef 22d agoSo you're saying the information is reliable?
- villgax 22d agoBasically the EvLeaks of AI lol
- g00afthrowaway 22d agoTheir negative reviews are more reliable than their positive reviews.
- amoss 22d agoSome of it is well sourced but the remainder is well sauced.
- neevans 22d ago[dead]
- theideaofcoffee 21d agoSo this is like the government attacking journalists for surfacing damning information, when in reality it's the government doing bad shit? Don't go after the interns doing the leaking, it's just easier to malign the messenger.
- iFire 22d agoSo out of all the inference only chips which ones can I buy? The only report is a smartnic fpga from Alibaba where we take an onnx design and write our own. https://essenceia.github.io/projects/alibaba_cloud_fpga/ https://essenceia.github.io/projects/alibaba_cloud_fpga/ On my M2 Pro Mac Mini the ANE only allows 2 gigabytes compared to the Metal GPU which can use the system ram. Currently playing with https://www.asus.com/motherboards-components/ai-accelerator/ugen/ugen300-usb-8g/ https://www.asus.com/motherboards-components/ai-accelerator/... which is a 4bit, 8bit and 16 bit ai inference chip with 8 gigabytes of ram. The UGen300 has the Hailo-10H chipset. The ASUS Store price for the ugen300-usb-8g costs $365.00 Canadian dollars.
- Traster 22d agoThis semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening in the industry and it doesn't seem like you can trust this as far as you can throw it.
- automatic6131 22d agoSemianalysis is an AI hype organisation, not a serious, unbiased semiconductor reviewer/journalist like chipsandcheese nor a documentarian of the semiconductor industry like Asianonmetry (as it relates so strongly to the modern economies of Asia). If you see something from semianalysis, you can simply ignore it.
- hermitShell 21d agoI tend to agree, but the Semianalysis + Dwarkesh side of reporting still does surface interesting information. You just have to take it all with a grain of salt, as indeed it is more hype focused, and look for real information hidden in the noise. And possibly to be a bit more entertained as you do so.
- automatic6131 21d agoI suppose access journalism does have to get something out of the bargain, however small it's not zero. If that's the way you like to spend idle time, well it takes all kinds. I don't understand competitive scrabble either.
- desterothx 19d agoThe trouble is the article is so poorly written I don't want to look for hidden information, i want to close out. I'm guessing ai wrote it, the information is constantly repeating, and not even always consistently
- 21d ago
- blt 22d agoIt can't come soon enough that AI workloads get their own specialized hardware to free up the general-purpose devices for general-purpose usage again.
- gbraad 22d ago... better than Blackwell in this specific case; which can also lead to OpenAI creating models that will only work on their own chips; vendor lock-in. When will they rename themselves?
- dev1ycan 22d ago"human speaks at... tokens" Yeah, stopped reading there, this is obviously some deranged sam altman paid blog post, I can't wait for the bubble to pop just so his newly launched chip falls flat on his face.
- philipwhiuk 22d agoIn the near term , the real question isn't whether it's better - it's whether they stop buying as many NVIDIA chips as they can.
- aurareturn 22d agoOne of the major advantages of Nvidia GPUs is that they can do both training and inference. If you make an inference only chip, you better be damn sure that it's significantly better than Nvidia's GPUs at it. Otherwise, it's better to buy Nvidia' GPUs because they're more flexible. You can do a big training run, then use them for inference right after.
- shelldon42 22d ago[dead]
- crate_88 22d ago[dead]
- kumarski 22d agolamb-labs.com built some cool stuff.... I'm inclined to believe we're going to see static model chips everywhere, but I'm no expert in software. I was early at efabless.com well now chipfoundry.io - they've done about 800 chip tape outs. They've been doing open source silicon tape outs for a decade plus. Founder recently built this: https://nativechips.ai https://nativechips.ai --- not involved but I'm inclined to believe it's the future of where the market is going. I'm skeptical of many of the AI chip design startups and whether they've actually taped out chips and how many and at what scale.
- redlewel 21d agomassive aura loss from a moronic name though
- kaveh_h 21d agoNVIDIA acquired Groq and have already integrated their technology in their solution. It seems it can really increase inference performance per energy used. I don’t believe The benchmark compares Jalapenjo together with this solution. https://nvidianews.nvidia.com/news/nvidia-groq-3-lpx-now-in-full-production-with-world-class-speed-for-agentic-ai https://nvidianews.nvidia.com/news/nvidia-groq-3-lpx-now-in-...
- fancyfredbot 18d agoWhat Jalapeño has done is pair huge bandwidth (almost as much as a full nvidia rubin) with a less powerful processing core. This means that even during the inference stage, when most architectures are bandwidth bound, they are compute bound instead. This is why they aren't using multi token prediction, as MTP only helps you when you are bandwidth bound. The power saving they get is likely from a much lower clock.