6 ms·
I don’t see how you could run Qwen3.8 27B on 16GB of memory that’s shared with Linux. Are people running models at 2bit quants? Are they even worth bothering wi
by teaearlgraycold 11d ago
I don’t see how you could run Qwen3.8 27B on 16GB of memory that’s shared with Linux. Are people running models at 2bit quants? Are they even worth bothering with? I had assumed you go down to 4bit and if you need to go smaller you have to lose parameters.
- beacon294 11d agoYou can, it's kind of cool to have this capability on something gaming at such a low price tag. Its not optimal for your time but beats nothing by a LOT. And that model is pretty reliable. https://unsloth.ai/docs/basics/dynamic-3.0-ggufs https://unsloth.ai/docs/basics/dynamic-3.0-ggufs
- carlos_rpn 8d agoYou can run a 4bit Qwen3.8 27B on a 8GB GPU, though on my RX 6650 XT it will only reach about 3 tokens/s. If you can give the GPU units enough RAM to load the whole thing, it may be a little faster. I'm not sure why would you want to run something at 3-4 tokens/s though. At least try a MoE model. I get about 30 tokens/s from those.