5 ms·
It can be both. A mistake in AD primitives can lead to theoretically incorrect derivatives. With the system I use I have run into a few scenarios where edge cas
by goosedragons 1y ago
It can be both. A mistake in AD primitives can lead to theoretically incorrect derivatives. With the system I use I have run into a few scenarios where edge cases aren't totally covered leading to the wrong result.
I have also run into numerical instability too.
- froobius 1y ago> A mistake in AD primitives can lead to theoretically incorrect derivatives Ok but that's true of any program. A mistake in the implementation of the program can lead to mistakes in the result of the program...
- goosedragons 1y agoThat's true! But it's also true that any program dealing with floats can run into numerical instability if care isn't taken to avoid it, no? It's also not necessarily immediately obvious that the derivatives ARE wrong if the implementation is wrong.
- srean 1y ago> It's also not necessarily immediately obvious that the derivatives ARE wrong if the implementation is wrong. It's neither full proof or fool proof but an absolute must is a check that the loss function is reducing. It quickly detects a common error that the sign came out wrong in my gradient call. Part of good practice one learns in grad school.
- froobius 1y agoYou can pretty concretely and easily check that the AD primatives are correct by comparing them to numerical differentiation.
- godelski 1y agoI haven't watched the video but the text says they're getting like 60+% error on simple linear ODEs which is pretty problematic. You're right, but the scale of the problem seems to be the issue