8 ms·
Just a couple days ago I learned about ninfer (https://github.com/Neroued/ninfer https://github.com/Neroued/ninfer) and on RTX 5090 I can now get ~200 tok/s and
by pllbnk 13d ago
Just a couple days ago I learned about ninfer (https://github.com/Neroued/ninfer https://github.com/Neroued/ninfer) and on RTX 5090 I can now get ~200 tok/s and over 400 tok/s on concurrent requests which is plenty fast for a local model of this strength.
- beastman82 13d agocan't second ninfer enough. amazing tech
- lowbloodsugar 13d agoOk, I need to try that. I'm getting 45tok/s with vLLM on my 6000. >600tok/s concurrent, but 45tok/s single request.
- pllbnk 13d agoEven without ninfer I would get over 80 on LM studio with default settings, so it should be noticeably more on 6000. You might want to try different a different inference engine or settings.
- jakswa 13d agodang only for certain nvidia GPUs, had my hopes up
- aizk 12d agoIs there an equivalent but for 4090s?
- pllbnk 12d agoThe repository has many forks, suggesting that folks are trying to (vibe) code support for different GPUs. Might be worth a shot.