29 ms·
It's a noble cause, but there are probably bigger levers to pull than the model size if you care about environmental impact. If you're running Qwen3.8-27B on e
by c7b 1mo ago
It's a noble cause, but there are probably bigger levers to pull than the model size if you care about environmental impact.
If you're running Qwen3.8-27B on energy-efficient hardware like a Mac or a DGX Spark instead of an API (likely running on H100s), I'm sure you're having much more of an impact than you would by switching to, say, a 9B coding-only model on the same hardware. The thing is, I think you won't be able to go orders of magnitude smaller, because a lot of the usefulness of LLMs comes from emergent smartness, and you typically need a minimum amount of complexity to see such emergent phenomena (and I think we're pretty far from understanding this kind of emergence, much further than from the next model generation that annihilates the current one on benchmarks yet again).