6 ms·
No it isn't. Best and worst and ill-defined anyway but the chess ELO score of various LLMs has fluctuated up and down, it's not been montonically increasing. Wh
by fragmede 4d ago
No it isn't. Best and worst and ill-defined anyway but the chess ELO score of various LLMs has fluctuated up and down, it's not been montonically increasing. What is the best answer to "how do I make cocaine"? The models are getting larger, with more compute and RAM backing them, but that doesn't automatically make them better if you don't define how you're measuring better-ness.
- famouswaffles 4d agoNone of the frontier labs care about Chess as it's already a solved problem. If they did, the models would be much better. It's really not that hard. Google has a paper on grandmaster level chess without search from transformers. Better obviously means better, like how they became better than they were 6 months and a year ago.
- vdomi 4d agoI would describe better as how much of my work I can delegate to the agent. Right now I'm delegating much more to Astra high than 6 months ago to Opus 4.6. Every dev has this feeling, it's weird to even argue what a better model/harness means.
- deleted 4d ago[deleted]
- deleted 4d ago[deleted]
- drxzcl 4d agoChess is not solved in any meaningful sense of the term. Computers have been better than humans since the 90s, but better chess programs are released all the time.
- famouswaffles 3d agoIt's solved in that we have had grossly superhuman capabilities for some time. It's not interesting for frontier labs. I suspect you understand this and the greater point so why be needlessly pedantic ?
- qlte 4d ago"Better" is not one dimensional across all use cases even if model capabilities are improving in aggregate. e.g. If someone said "this is the worst they'll ever be" in response to some writing with obvious LLM cliches in 2024, I'm not convinced that prediction was actually correct. The focus of OpenAI/Anthropic pivoted aggressively to the agentic performance arms race instead of making a more human sounding chatbot so regressions in writing ability aren't really a concern anymore if agentic benchmarks improve. The first time I heard a recommendation to use Claude was specifically because it sounded much more "human" and natural than ChatGPT. Fast forward to now and idiosyncratic Claude-isms repeated every other sentence and its convoluted verbosity has become a widely mocked meme.
- famouswaffles 3d agoRight but we're talking about a single subject here - mathematics that labs are incentivized to keep improving for some time. >e.g. If someone said "this is the worst they'll ever be" in response to some writing with obvious LLM cliches in 2024, I'm not convinced that prediction was actually correct. 2024 creative writing prose was...the last few versions have stalled, but I think they're still better than 2024.