7 ms·
Frontier models went from not being able to count the number of 'r's in "strawberry" to getting gold at IMO in under 2 years [0], and people keep repeating the
by charleshn 1y ago
Frontier models went from not being able to count the number of 'r's in "strawberry" to getting gold at IMO in under 2 years [0], and people keep repeating the same clichés such as "LLMs can't reason" or "they're just next token predictors".
At this point, I think it can only be explained by ignorance, bad faith, or fear of becoming irrelevant.
[0] https://x.com/alexwei_/status/1946477742855532918 https://x.com/alexwei_/status/1946477742855532918
- bwfan123 1y ago> At this point, I think it can only be explained by ignorance, bad faith, or fear of becoming irrelevant. Based on the past history with frontier-math & AIME 2025 [1],[2] I would not trust announcements which cant be independently verified. I am excited to try it out though. Also, the performance of LLMs was not even bronze [3]. Finally, this article shows that LLMs were just mostly bluffing [4]. [1] https://www.reddit.com/r/slatestarcodex/comments/1i53ih7/frontiermath_was_funded_by_openai_and_they_have/ https://www.reddit.com/r/slatestarcodex/comments/1i53ih7/fro... [2] https://x.com/DimitrisPapail/status/1888325914603516214 https://x.com/DimitrisPapail/status/1888325914603516214 [3] https://matharena.ai/imo/ https://matharena.ai/imo/ [4] https://arxiv.org/pdf/2503.21934 https://arxiv.org/pdf/2503.21934
- yeasku 1y agoOpen AI is 10 years old and and llm just told me a dolar is 1.03 euros.