6 ms·
On the DGX I get 44.5 tokens per second (NVFP4). With 8 concurrent it's 241 t/s total. I am using the PrismaAQUA standard 9.7 t/s + Dflash2 30 t/s + torch-c
by decide1000 24d ago
On the DGX I get 44.5 tokens per second (NVFP4). With 8 concurrent it's 241 t/s total.
I am using the PrismaAQUA
standard 9.7 t/s
+ Dflash2 30 t/s
+ torch-compile 37 t/s
c8 = 177 t/s
- SwellJoe 24d agoWhat model? Also, I don't know what "the PrismaAQUA" means, ddg thinks it's a CPAP machine, which seems unlikely to help with inference performance. Also, 4-bit has measurable intelligence loss. Sometimes worth it, but, at this size models are barely smart enough at 8 or 6.
- decide1000 24d agoQwen3.8-27B-PrismaAQUA-5.5bit-vllm The output quality is higher. It's held at full precision (not quantized).