7 ms·
What is "sticky" in this context?
by greazy 8d ago
What is "sticky" in this context?
- robrenaud 8d agoExperts vary per token in MoE, there is maximum flexibility. Good for driving down loss, bad for locality/gpu memory/bandwidth. If expert selection were more constrained, inference systems could take advantage of it. Keeping experts cached would mean not needing to load them from disk/ram every token.