4 ms·
It depends on your tolerance level for having less than frontier LLM capabilities. First level worth trying, Qwen 3.6 35B A3B with a 16 GB VRAM (example 1x 5060
by jononor 1mo ago
It depends on your tolerance level for having less than frontier LLM capabilities. First level worth trying, Qwen 3.6 35B A3B with a 16 GB VRAM (example 1x 5060ti 16gb, 600 USD for the card) with partial GPU offloading. Next level would be Qwen 3.6 27B / new Muse Spark / Gemma 31B with 32 GB VRAM (2x 5060ti or 1x 9700 Pro). Third level would be DeepSeek V4 Flash with 192 GB VRAM (2x Strix Halo at some 8000 USD total). These models can be tried on OpenRuouter etc, or you can deploy vLLM on rented GPUs to get a feel for what level you would want before committing to buying hardware.