8 ms·
> I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months. .... :T Considering the 8B model uses 5
by x-complexity 29d ago
> I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months.
....
:T
Considering the 8B model uses 53 billion transistors, that's 6.625 transistors per parameter.
https://taalas.com/products/ https://taalas.com/products/
Assuming they can get it down to 3 (somehow), that's still 300 transistors, or 5.565 RX 9070s.
https://www.techpowerup.com/gpu-specs/radeon-rx-9070.c4250 https://www.techpowerup.com/gpu-specs/radeon-rx-9070.c4250
You're looking at
1) waiting for another 3-5 generations of transistor improvements before it can fit into a single conventional chip, or
2) another generation before getting a monster of a chip (1000+ mm^2), and prices for flawless etching scale quadraticly (likely $1000+ for manufacturing costs alone).
Could happen, but it's a long shot for a market that could be satiated by specialized accelerators.
- mdp2021 29d agoIn Taalas HC2 a chip embeds 20b parameters, and the declared idea is linking the chips. A card with two of them chips and you can already have a dense Qwen at staggering speeds.
- x-complexity 28d agoChiplet-style layouts could cut the etching quality requirements per chip down, but it still can't avoid the base cost for manufacturing silicon.
- mdp2021 28d ago> it still can't avoid the base cost for manufacturing silicon What costs are you talking about and why would them be a problem? If it is the price: «Kharya says it costs 100x as much to train a model then to get a customize HC chip in reasonable volumes from Taalas» ( https://www.nextplatform.com/compute/2026/02/19/taalas-etches-ai-models-onto-transistors-to-rocket-boost-inference/4092140 https://www.nextplatform.com/compute/2026/02/19/taalas-etche... ). That is thinking about an LLM (logical) producer and server. For mass production, the costs go down. And in the case of a ~100b model as the poster mentioned, they would be just single PCI cards with 5 or 6 HC2 chips: doable and practical. I see more potential problems in the positioning of the SRAM - but not a real problem given that excellent team. To get a proper idea of the costs the architecture of the HC2 will have to be clearer.