6 ms·
The FP8 version from DeepSeek themselves [0], around 1800 tps prefill and 45 tokens per second decode. I’ve been running a custom VLLM image with b12x as well
by redrove 2mo ago
The FP8 version from DeepSeek themselves [0], around 1800 tps prefill and 45 tokens per second decode.
I’ve been running a custom VLLM image with b12x as well as nvfp4_ds_mla.
I would say it’s quite fantastic in day to day, I use it mostly in Hermes and sometimes for coding.
I have qwen 3.6 27b on an rtx 6000 pro as well so I use that as a workhorse in pi with DS as a reviewer/planner.
[0] https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Edit: I think you may have misread my post. k3s is NOT kimi k3, and I did mention I was running deepseek.
- pixelesque 2mo agoThanks - yeah, sorry, I mis-read that as you using both DS and K3...
- reinitctxoffset 2mo agoI'm doing K3 with SoL kernels, zero hassle, myelin to sweep up the crap, on shot deploy to vast or runpods. Free to the community. I have to do like, low paying web dev to fund it, so I can't promise timelines, but it's coming and it will be free to anyone and fast as fuck. I estimate about 20 bucks an hour at current spot rates in the hundreds if not thousands of tokens per second.
- reinitctxoffset 2mo ago@pg @dang you're asking me how a watch works. let's just try to keep an eye on the time. federal felony prosecution.
- redrove 2mo agoNo worries! Would love to run K3 but lack the hardware for a 2.8T model, like most people.