8 ms·
Not suggesting a provider but if you're willing to get a Chinamod GPU, go get RTX 3080 20G or RTX 2080 Ti 22G. Get a couple of them and you can probably run Qwe
by jan_Sate 19d ago
Not suggesting a provider but if you're willing to get a Chinamod GPU, go get RTX 3080 20G or RTX 2080 Ti 22G. Get a couple of them and you can probably run Qwen3.8-27B at a reasonable speed.
I've got a RTX 3080 20G for $450 like a year ago. With llama.cpp, bf16 kv cache, kv cache offloaded to RAM, Qwen3.8 27B UD-Q4_K_XL, single RTX 3080 20G, I got 10 tok/s initially and it dropped to 5 tok/s at 50k context. I'm thinking of getting another Chinamod GPU.
Building a dual Chinamod GPU machine would only cost like $1000~$1500. It's not too bad compared with the alternatives!
- cyanydeez 19d agoA year ago was prior to the memory cartel. You're probably off by 2x.
- jan_Sate 19d agoI've checked. I can still get a 3080 20G for $525, or 2080 ti 22G for $380. The price has increased but not too bad. Probably the modding had prevented it from gaining too much price increase.
- BoredomIsFun 19d ago> single RTX 3080 20G, I got 10 tok/s initially Add $200, buy a used 3060 and you won't need to offload cache to RAM. Yo'd have like 50-60 t/s with MTP enabled.
- jan_Sate 19d agoaww. I used to have a 3060 and I upgraded to this 3080 20G. Maybe it's time for me to get a 3060 back again. :P
- gstar 19d agoOr, theres the option of a 32gb v100 (about $600USD on taobao etc), where you can get 1200 of prefill at 80 of decode: https://github.com/geoffwatts/ninfer-v100 https://github.com/geoffwatts/ninfer-v100 - that's _really_ cheap inference, and it's not a modified card - you just need to add a blower or water block.
- slim 18d agoI run qwen3.8 on 5060ti 16G RAM. If you can do with Q3 and 64k of context. It works great at 25t/s