7 ms·
They're running like 10 tokens per second, constantly timing out, and launching new products? How about deploy some inference GPUs first
by bofadeez 2mo ago
They're running like 10 tokens per second, constantly timing out, and launching new products?
How about deploy some inference GPUs first