11 ms·
Completely agree in principle, I'd expect this when minimizing entropy over any text incl. code. However, evals across variety of domains show that LLMs can rea
by _false 1y ago
Completely agree in principle, I'd expect this when minimizing entropy over any text incl. code. However, evals across variety of domains show that LLMs can reach (and even surpass) expert performance[^1].
[1]: https://arxiv.org/abs/2508.17669 https://arxiv.org/abs/2508.17669