6 ms·
Imagine this, and sucesor models on Cerebras or other silicon...
by manofmanysmiles 1mo ago
Imagine this, and sucesor models on Cerebras or other silicon...
- WASDx 1mo agoThat might actually compensate for the overthinking, if it can think really fast. Dense models are easier than MoE to put on silicon. https://chatjimmy.ai/ https://chatjimmy.ai/ is getting 16k tps with an 8B model. Extrapolating that gives nearly 5k tps for 27B. And we're still early in this technology. If tps is so high, a compaction step could be performed over every thinking turn to keep context size down.
- Moduke 1mo agoVery exciting indeed. It is in the works. Their current dense offering, Gemma 4 31B, sits at ~1800t/s https://news.ycombinator.com/item?id=49308715 https://news.ycombinator.com/item?id=49308715