7 ms·
Doesn't matter if the chip is 100x-1000x more efficient and faster, and you can just make a new one for new weights. Imagine being able to run GPT Sol at 1k tok
by WithinReason 22d ago
Doesn't matter if the chip is 100x-1000x more efficient and faster, and you can just make a new one for new weights. Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Or Qwen 3.8 27B at 10k tokens/second. The super long thinking that makes qwen so effective would take a couple of seconds.
- ben_w 21d ago> Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that: A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU. There, inference costs declined from $15 per million tokens in May 2024 to $0.12 per million tokens by December 2024 (Phi 4). - https://hai.stanford.edu/assets/files/hai_ai-index-report-2025_chapter1_final.pdf https://hai.stanford.edu/assets/files/hai_ai-index-report-20... 15/0.12 -> factor of 125 cost reduction in 7 months. But that may well be an extreme case. To show how broad the range is, another quote from the same publication: Depending on the task, LLM inference prices have fallen anywhere from 9 to 900 times per year.
- WithinReason 21d agoThis is from 2025, how about the last 6 months?
- ben_w 21d agoYou tell me. Most of these reports take that long to get published, or even longer. Sometimes I even see new-ish reports talking about 4o.
- jurgenburgen 20d agoI’ll admit that it would be an interesting product, assuming you could buy the chips / cards and slot it in commodity hardware.