5 ms·
It's larger than previous V4 Flash. 552B in ~FP4, 306GB. 196B of FP8 Engrams, another 204GB, not necessary to keep in RAM. KV cache sees another 4x
by petu 7d ago
It's larger than previous V4 Flash.
552B in ~FP4, 306GB.
196B of FP8 Engrams, another 204GB, not necessary to keep in RAM.
KV cache sees another 4x size reduction, just 900MB for 1M.
So 384GB needed for a chance of achieving useful speeds. Three Sparks or quad RTX PRO 6000.
- npodbielski 7d agoOr two gorgon halos?
- Tuna-Fish 6d agoOr one Medusa Halo.
- hypfer 6d agoQuestion is how many of those experts one needs to keep in vram for a given workload. I could imagine (though I might be _very_ wrong there) that for example coding does not live in all of them. Maybe 1/3? Do we have real numbers there? So maybe one can get away without much performance penalty by doing some LRU stuff?