6 ms·
Macs have excellent generation speed, and the new Ultra will positively smash that at 1.2TB/sec of bandwidth. For example, that new 176B parameter Qwen model wo
by anon373839 16d ago
Macs have excellent generation speed, and the new Ultra will positively smash that at 1.2TB/sec of bandwidth. For example, that new 176B parameter Qwen model would generate tokens at ~200 tokens/sec.
Macs don’t have very good prefill, though. So it’s important to use a model serving stack that has excellent prompt caching and use a harness that won’t bust the cache.
I’m cross-shopping DGX Sparks and M5 Studios, and having a hard time deciding because they have exactly opposite characteristics for prefill and decode.