6 ms·
There are benchmarks that humans score close to zero on average and the top LLM scores 25%. https://epoch.ai/frontiermath/the-benchmark https://epoch.ai/fronti
by bufferoverflow 2y ago
There are benchmarks that humans score close to zero on average and the top LLM scores 25%.
https://epoch.ai/frontiermath/the-benchmark https://epoch.ai/frontiermath/the-benchmark
- pama 2y agoIf anyone from epoch.ai is reading this, it would be nice to link the toplevel result for o3 to this page.