5 ms·
This initial round of benchmarking was to understand if there was any usecase here at all and I think there is. In a follow up, I'll be trying to answer questio
by eso_logic 2mo ago
This initial round of benchmarking was to understand if there was any usecase here at all and I think there is. In a follow up, I'll be trying to answer questions like this. How big of a model can you fit on 4x M60, 4x P100, 4x V100? What are the tok/second when varying context length?
Do you have a set of models you'd like me to look at?
- russianGuy83829 2mo agoThat's great. Personally, I'd interested in Qwen3.6-27B and deepseek V4 flash (or pro), with contexts above 60k. They seem to be popular and have good coding performance. I'd appreciate numbers on a single or two GPUs where a quantized version fits reasonably into the VRAM (Qwen in 16 or 24GB). 4 older GPUs approach a used 3090 in price, and the 3090 has better support for speedups like MTP. So cheaper but slower looks like a reasonable target to me.
- eso_logic 2mo agoNo problem. Varying context size is a common request I've been getting as well. Personally I'm looking forward to seeing how much we can cram into the ancient K80's 24GB of VRAM :0
- russianGuy83829 2mo agoThank you, looking forward to it. I just saw this simple patch to enable MTP (potentially 2x performance) on older GPUs (Kepler etc), so maybe it will work for you https://github.com/ggml-org/llama.cpp/pull/25680 https://github.com/ggml-org/llama.cpp/pull/25680 Also, for Qwen, the 4 bit _XL quantization seems to have a good balance of performance to size.
- NortySpock 2mo agoSimilar interest here, possibly including if qwen 3.6, Gemma4 or DiffusionGemma (with the largest quants that will fit in a single card) will offer, say, 50 tokens-per-second (fast enough for interactive human-in-the-loop code research, print-f iterations on code to debug things, etc; or let the LLM churn on a problem for a minute while I step out to handle something else), context of up to 200k preferred. Also if nothing else the below project lets you use an NVidia graphics card as low-latency swap, which has been nice as a buffer as RAM prices remain high and leaves me eyeing that 24GB card you mentioned as an alternative... https://github.com/c0deJedi/nbd-vram https://github.com/c0deJedi/nbd-vram