6 ms·
true, but no reason the predictor model couldn't use linear attention (i.e. mamba, GDN etc) to predict KV caches
by somnial 3mo ago
true, but no reason the predictor model couldn't use linear attention (i.e. mamba, GDN etc) to predict KV caches