7 ms·
This confused me at first as well.. inactive experts skip compute, but weights are sill loaded. So memory does not shrink at all. I found this visualisation he
by functional_dev 5mo ago
This confused me at first as well.. inactive experts skip compute, but weights are sill loaded. So memory does not shrink at all.
I found this visualisation helpful - https://vectree.io/c/sparse-activation-patterns-and-memory-efficiency-in-zero-activated-weights https://vectree.io/c/sparse-activation-patterns-and-memory-e...