6 ms·
So, what's the most affordable way for a pleb who doesn't own 17 H100s to use Kimi K3 or Qwen 3.8?
by joey64 2mo ago
So, what's the most affordable way for a pleb who doesn't own 17 H100s to use Kimi K3 or Qwen 3.8?
- vanillax 2mo agoyou cant. The best you can do is Qwen 3.6 27b with a 24gig ( or cumaltive gpus ) to get to 24gb vram. ala 3090, mac with 36gb ram, amd cards, halo strix amd, dgx spark etc. Lots of youtube videos out there.
- Alpha3031 2mo agoWell, if you're happy with around (as in within an order of magnitude or two of) 0.1 tokens per second... I believe that's around what people are getting when loading MoE weights from NVMe.
- svachalek 2mo agoYou've got to consider how much power that uses though. Depending where you live, some of these providers can serve it for less than you pay for power.
- drnick1 2mo agoThere will be smaller versions in the 10-30B parameters range that can run on consumer GPUs.
- svachalek 2mo agoI haven't seen either of these running outside their creator's services yet, but typically you can watch services like openrouter or nano-gpt for it to show up at a (usually small) discount.