5 ms·
An LLM is a set of structured matrix multiplies and function applications. The only potentially non-deterministic step is selecting the next token from the fina
by philipswood 5mo ago
An LLM is a set of structured matrix multiplies and function applications. The only potentially non-deterministic step is selecting the next token from the final output and that can be done deterministically.
- jmalicki 5mo agoMatrix multiplication on GPUs is non-deterministic. As are things like cumsum() https://docs.pytorch.org/docs/2.11/generated/torch.use_deterministic_algorithms.html#torch.use_deterministic_algorithms https://docs.pytorch.org/docs/2.11/generated/torch.use_deter... This comes down to map reduce and floating point's lack of associativity. You see the same thing with OpenMP on CPUs. People are constantly claiming determinism in LLMs that is just not there.
- vrighter 5mo agowell just run all inference on the cpu, single threaded /s
- zadikian 5mo agoEven if it were reproducible, realistically most people are using some service like Claude that makes no guarantee that the model or hardware didn't change. Which is fine, it doesn't need reproducibility. This is interesting though, I didn't know PyTorch had a debug mode for reproducibility.
- jmalicki 5mo agoEven with this debug mode, a different batch size can give different results for the same input - e.g. your tensor multiplies might use different blocking, hence different associativity. I posted that to show that at a bare minimum, there is some pretty extreme nondeterminism (though probably mild in effect) in even the most pedestrian GPU inference, unless you go to the extreme of using the debug mode and taking the potential performance hit.