11 ms·
Digging into the low rank structure of the gradients, instead of the weights seems like a promising direction for training from scratch with less memory require
by hantusk 3y ago
Digging into the low rank structure of the gradients, instead of the weights seems like a promising direction for training from scratch with less memory requirements: https://twitter.com/AnimaAnandkumar/status/1765613815146893348 https://twitter.com/AnimaAnandkumar/status/17656138151468933...
- hantusk 3y agoSimo linked some older papers with this same idea: https://twitter.com/cloneofsimo/status/1765796493955674286 https://twitter.com/cloneofsimo/status/1765796493955674286