6 ms·
That statement gives a lot of insights into the possible model size. Llama 3.1 405b runs @ ~900t/s: https://www.cerebras.ai/blog/llama-405b-inference https://w
by _boffin_ 3mo ago
That statement gives a lot of insights into the possible model size.
Llama 3.1 405b runs @ ~900t/s: https://www.cerebras.ai/blog/llama-405b-inference https://www.cerebras.ai/blog/llama-405b-inference
From initial vibe research (which is totally not correct by any means) ~13.5k concurrent streaming clients capacity