4 ms·
> We have a CiM with Taalas HC1 My understanding is that Taalas HC1 is "mask-ROM" fabricated at 6 nm for bulk model base-weight storage, and some SRAM for KV c
by topspin 1mo ago
> We have a CiM with Taalas HC1
My understanding is that Taalas HC1 is "mask-ROM" fabricated at 6 nm for bulk model base-weight storage, and some SRAM for KV cache and other bits:
https://www.eetimes.com/taalas-specializes-to-extremes-for-extraordinary-token-speed/ https://www.eetimes.com/taalas-specializes-to-extremes-for-e...
"On the HC1, the model and its weights are stored on the chip using a mask-ROM-based recall fabric paired with a (programmable) SRAM"
I don't believe that's CiM as you advocate.
> I believe that "practical" as in "replaceable" is also a fundamental property of what we desire in this field
I suspect that there is a important frequency factor in in the "replaceable" calculus. Already I see people dragging their feet about adopting newer models once they've found familiarity with some older model: "good enough" is a thing. I know there are industries where "validated" is a concept, and they do not ride wave crests. So, if we imagine that as all this eventually shakes out and we're not replacing models every few months, but instead with about the same frequency as our cell phones or similar, the ROM model works. If the performance and price make this pattern highly appealing, then that's what will win, certainly for local inference. If some datacenter operator could, today, adopt a ROM approach that cut their power budget by a large factor, but had to suffer 2-3x longer model update cycles, they'd likely consider it.
For better or worse.
I have no problem with CiM as a concept. If it can reduce power/size/cost then it's another avenue that inference will probably incentivize, where incentive has previously been insufficient. As we both agreed long ago in this thread this new era is motivating things that were previously neglected, and CiM is possibly a part of that. My dream is that all of these get a hard look as people try to figure out how to run all of this without enormous gigawatt sucking datacenters that rival DOD program budgets.
- mdp2021 1mo ago> I don't believe that's CiM as you advocate You missed the whole point of Taalas HC1: that it is Compute-in-Memory. > 2. Merging storage and computation // Modern inference hardware is constrained by an artificial divide: memory on one side, compute on the other, operating at fundamentally different speeds. // This separation arises from a longstanding paradox. DRAM is far denser, and therefore cheaper, than the types of memory compatible with standard chip processes. However, accessing off-chip DRAM is thousands of times slower than on-chip memory. Conversely, compute chips cannot be built using DRAM processes. // This divide underpins much of the complexity in modern inference hardware, creating the need for advanced packaging, HBM stacks, massive I/O bandwidth, soaring per-chip power consumption, and liquid cooling. // Taalas eliminates this boundary. By unifying storage and compute on a single chip, at DRAM-level density, our architecture far surpasses what was previously possible. https://taalas.com/the-path-to-ubiquitous-ai/ https://taalas.com/the-path-to-ubiquitous-ai/