5 ms·
>The actual benchmark improvements are marginal at best GPT-5 demonstrates exponential growth in task completion times: https://metr.org/blog/2025-03-19-measu
by z7 1y ago
>The actual benchmark improvements are marginal at best
GPT-5 demonstrates exponential growth in task completion times:
https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com...
- hk__2 1y agoWhat do you mean? A single data point cannot be exponential. What the blog post say is that the ability to solve tasks of all LLMs is exponential over time, and GPT-5 fits in that curve.
- z7 1y agoYes, but the jump in performance from o3 is well beyond marginal while also fitting an exponential trend, which undermines the parent's claim on two counts.
- adammarples 1y agoActually a single data point fits a huge range of exponential functions.
- usaar333 1y agoNo it doesn't. If it were even linear compared to o1 -> o3, we'd be at 2.43 hours. Instead we're only at 2.29. Exponential would be at 3.6 hours