5 ms·
Last statement is a bit sus... Muon computes matrix sign function which can be defined as setting singular values to 1, though you can also define it without SV
by big-chungus4 2mo ago
Last statement is a bit sus... Muon computes matrix sign function which can be defined as setting singular values to 1, though you can also define it without SVD. Muon itself doesn't use SVD because it uses a faster method to compute matrix sign. Adam doesn't do anything related to SVD or singular values. Also not sure what you meant by "second order singular values"
- jmalicki 2mo agoADAM is related if your second derivative matrix happens to be diagonal. Of course, it takes about 5 minutes to show that any DNN is going to have very very high magnitude off-diagonal terms by the way it's constructed, so pretending that a diagonal approximation is close enough is crazy.
- big-chungus4 2mo agoAdam doesn't use the second derivatives matrix, it uses second moments of the gradient, which is the diagonal of the uncentered covariance matrix, but neither of them are directly related to SVD or singular values anyway. There is a slight connection where Adam approximates full-matrix Adagrad which computes inverse square root of the convariance matrix, which you usually do using eigendecomposition, but on the covariance matrix SVD and eigendecomposition are equivalent (can easily be converted to each other), so you could use SVD to compute the inverse square root.
- jmalicki 2mo agoThe second moments of the gradient and the Hessian are absolutely related! See the Fisher Information, and the Cramer-Rao Lower Bound (an inequality on how much the inverse covariance matrix and the Hessian can differ). https://en.wikipedia.org/wiki/Fisher_information https://en.wikipedia.org/wiki/Fisher_information
- jmalicki 2mo agoI didn't find a direct proof earlier, just assertions - I've only seen it proved in textbooks that aren't linkable. Theorem 1, section 1.3, page 2 shows that the expected variance of the gradient of the loss function and the expected second derivative of the loss function are equal at the minimum. I hate that the ADAM paper did not talk about this, this is something that is hammered into anyone who has taken a mathematical statistics course. This has been an established fact in statistics for well over 100 years. https://courses.grainger.illinois.edu/ece563/fa2025/Notes11-IT2025.pdf https://courses.grainger.illinois.edu/ece563/fa2025/Notes11-... Away from the minimum they can diverge, but there is a close enough connection to make it an extremely useful approximation.