8 ms·
As a 5090 owner and local model enthusiast, I was hoping it would be 35B A3B so I could run it myself =(.
by hasteg 22d ago
As a 5090 owner and local model enthusiast, I was hoping it would be 35B A3B so I could run it myself =(.
- cpburns2009 22d agoYou can run the 27B released last week. I haven't tried it yet myself but the 3.6 version runs great on my 5090.
- Philpax 22d agoStrongly recommend https://github.com/Neroued/ninfer https://github.com/Neroued/ninfer, which can pull ~180 TPS on 5090 with 3.8, and 500 (!) with 3.6 35B-A3B.
- cpburns2009 22d agoI've been waiting for the dust to settle on this model so I can find a good runtime setup. I'm definitely bookmarking this. Thanks!
- hasteg 21d agoI've been running 27B a lot, I am honestly shocked at how well it performs. It's mind blowing how well the small qwen models (and particular, the 3.8 model) runs locally. It can create some really impressive toy coding projects. I haven't really done too much integrating into my actual workflows because I pay $100 for Claude Max, but I can see it being pretty helpful in that.
- Tuna-Fish 22d agoThe 27B one is great on a 5090. This one is basically aimed at macs, Strix halo and DGX Spark.
- latentsea 22d agoDepends on the rest of your hardware, and how the the n-gram weights work and if they can be streamed from SSD. If they can and you have 64GB system ram then you should actually be able to run it.