7 ms·
Out of curiosity, what's currently the best model I can use locally?
by brettgo1 1mo ago
Out of curiosity, what's currently the best model I can use locally?
- daemonologist 1mo agoWith an unlimited budget, Kimi K3 (which is quite comparable to this Qwen Max imo). With a normal budget/a PC you might already have, probably Qwen 3.6 27B.
- arjie 1mo ago$500k - Kimi K3 (maybe $250k? Haven’t done this one) $25k - DSv4 Flash $4k - Qwen 3.6 35A3B Q5 $1k - Qwen 3.6 27B Q4 Some people prefer the sense over the MoE YMMV.
- apitman 1mo agoThese numbers look about right based on my experiences as well. Though for a single user I think 2x DGX Spark (~$10k) runs DSv4 Flash fairly well right?
- colingauvin 1mo ago16 DGX Sparks can run K3 at a reasonable TPS. So that's $64k. 2 DGX Sparks can run DS4 at 1 million context with 50 TPS so that's $8k. 1 A4500 can run 35A3B. Those are about $1200 new. 27B actually takes more hardware to run than 35B because attention is done differently I believe and therefore KV Cache takes a lot of space. It will run on an A4500 but it's slow and context will be like 32k.