8 ms·
I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.
by zxspectrum1982 1mo ago
I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.
- Gigachad 1mo agoIt costs something like $300,000 for the hardware to run a model of that size. You'd pay that for a single model for 1-2 years? Not even the AI companies can justify that kind of spend which is why they keep extending the expected lifespan on their hardware in the accounting.
- lsaferite 1mo agoGiven the fact that Taalas was claiming 1/10th the hardware cost and 1/10th the power consumption, yeah, companies would absolutely jump at that. Now, whether they can achieve that in practice is yet to be seen.
- zxspectrum1982 1mo agoI'm expecting the Taalas MSIC version to cost a fraction of that. Then probably have some kind of cheap subscription to Anthropic for updates (yes, Taalas chips can receive a certain kind of updates: they have a small SRAM).
- mdp2021 1mo ago> It costs something like $300,000 for the hardware to run a model of that size You did not compute that as the cost for a speculative card from Taalas, right?
- Gigachad 1mo agoIt's the cost of the current nvidia hardware used to run these models. Of course all bets are off if you are accounting for some future chip that doesn't exist yet which could cost less.
- subroutine 1mo agoBut Claude Opus 4.6 is not really practical. Taalas' process seems targeted for edge models. Their proof of concept model, for example, is a heavily quantized version of Llama 3.1 8B and even then they acknowledge their custom 3-bit/6-bit representation causes model quality degradation. Taalas is going to have a tough time putting a trillion-parameter model on one conventional die. Their HC1 die is already near the maximum size that conventional lithography can expose. They claim they could partition the model across many chips, but I'm not sure if they have tested this process or what it means for compute. The basic storage arithmetic is unforgiving: for a one trillion parameters model at four bits it will take 50–100 chips. To service a sizable customer base will take thousands of 100-chip fabs. That all said, I'm bullish on this technology, and look forward to seeing it evolve.
- vatsachak 1mo agoYeah. But this kinda feels like a bandaid. Eventually someone will have to solve compute in memory at scale.
- Iolaum 1mo agoA really fast qwen-3.6-27B type of model could be useful. With a specialized harness and this speed I 'd expect it to find many applications. Implementing a coding plan is the minimum I can think of.
- momojo 1mo agoI'm sure life would find a way. I'd love to see what kind of power-harnesses people have to come up with to steer 16k tps QPU's (Qwen Processing Units) productively.
- andix 1mo agoWith thousands of token per second output it would be an enormous waste of resources. Such chips are clearly made to process thousands of conversations simultaneously. Not necessarily in parallel. All LLM workflows are turn based right now, there are often seconds between turns until tool calls finish or users type the next message. If the LLM response only takes a few milliseconds, the chip can process hundreds of other requests until the first conversation becomes active again.
- NiloCK 1mo agoNot so long ago, I was good enough for many coding tasks. But I found that things can change in a hurry. Yes, a cheap and fast Opus4.6 can drive a lot of value in current context. But if we continue to craft bigger-and-bigger balls of mud, Opus 4.6 may end up hitting its conceptual ceiling and unable to contribute. Winding the clock back on your statement gives: > I'd gladly pay for a Claude Sonnet 3.5 in silicon and use it for 1-2 years. Man, I dunno.
- kennywinker 1mo agoAssuming moore's law like progress, which I'm 100% sure isn't going to happen - I think we're at the top of the S curve already. But assuming dramatically increased intelligence every year this is still the exact same position as anyone who bought a computer in the last 5 decades. Yet, people did very much buy computers.
- inigyou 1mo agoBut that's what people said 6 months ago about whichever model was current 6 months ago, but you hate that model now.