5 ms·
14B even at Q4 isn't realistic for coding on a single 12GB RTX 3060. Token speed is too slow. After all they are dense models. You aren't getting a good MoE mod
by rdos 6mo ago
14B even at Q4 isn't realistic for coding on a single 12GB RTX 3060. Token speed is too slow. After all they are dense models. You aren't getting a good MoE model under 30B. You can do OCR, STT, TTS really well and for LLMs, good use cases are classification, summarization and extraction with <10B models.
- suprjami 6mo agoDual 3060s run 24B Q6 and 32B Q4 at ~15 tok/sec. That's fast enough to be usable. Add a third one and you can run Qwen 3.5 27B Q6 with 128k ctx. For less than the price of a 3090.
- rdos 5mo agoSure, two 3060 can pull usable performance on an usable LLM, but a single one can't (yet). > 3x RTX 3060 less tgab the price of a 3090 Interesting, here it is around the same. 200-250€ for a used 12GB 3060 and 600-800 for a used 3090€.