5 ms·
How fast is it on consumer-level hardware?
by bityard 26d ago
How fast is it on consumer-level hardware?
- toebee 26d agoWe got a rtx 4090 handling around 10 concurrent requests at 50 ms TTFA after some config changes / adjustment as it doesn’t have FP8. So this 50 ms TTFA thing is very much possible on consumer hardware.
- bt1a 26d agoI'm going to see how this shakes out on my machine with a few 3090s. I see you all are leveraging some custom cuda kernels, so it may not work out of the box on Ampere (30xx) architecture yeah?
- toebee 26d agoYep, might need some changes.
- jubilanti 25d agoI can buy a used car for the price of a used RTX 4090 ($2500-$3000), I wouldn't consider it consumer hardware. Prosumer maybe. Almost no consumer needs 10 concurrent requests. How fast does this run on a 3060 or CPU/iGPU only like an Intel Iris or AMD Navi? Or is your priority more commercial cloud services instead of local self hosted?