7 ms·
Llama 405B 506 tokens/second on an H200
- moondistance 2y agoSignificant further optimizations. FP8!
- 7e 2y agoAnd this is why nobody submits MLPerf against NVIDIA.
- greenknight 2y agoIts weird, i looked up whether AMD has any benchmarks on the 405B for the MI300x, and came across this one -- https://dstack.ai/blog/amd-mi300x-inference-benchmark/#tokensec-and-ttft-per-rps https://dstack.ai/blog/amd-mi300x-inference-benchmark/#token... From my understanding, it can get up to around 2500 tokens/s? Both are 8x units (h200 and MI300x)
- EgoIncarnate 2y agonot "an H200", "In the table above, tensor parallelism is compared to pipeline parallelism with each across eight GPUs"
- FanaHOVA 2y agoTitle on HN is wrong. The article says GPUs and it's referring to one of their 8xH200 boxes.