10 ms·
> - Math completely fails in longer contexts Not sure what longer contexts we're talking about but didn't we have an old math problem optimized, which even the
by fakwandi_priv 2mo ago
> - Math completely fails in longer contexts
Not sure what longer contexts we're talking about but didn't we have an old math problem optimized, which even the LLM itself was surprised about, just a week ago? Something which wasn't possible 6 months ago.
- rf15 2mo agoI mean calculations, not mathematical proofs
- Barbing 2mo agoIf they use Python to fill the gap, and the end user doesn’t have to know or care, is it unfair to assess this as progress and attribute the progress to the _system_? OK, the core technology that is the language model still can’t math as well as you’d hope, but how about the end result users see from the system when they interface with it? “Did you know humans are better at flying today than they were a thousand years ago?” ‘No they’re not, they need planes.’ Technically correct in a way but isn’t it kind of annoying to be so stubbornly pedantic when the context is speed of reaching Point B from Point A?
- rf15 2mo agoYou are correct, the frameworks around it have improved. In that regard, my assessment is unfair: I only judge the underlying technology and what is sold by the sota providers, with the premise of what it's like when you start fresh. You can achieve a lot by coding around the issues, but that's kinda against the point of 'AI', is it?
- wahnfrieden 2mo agoNo, its capabilities with a harness are what we are interested in. Your assessment is only relevant to benchmarking, not practical value.
- Barbing 2mo agoorwin‘s response in a cousin comment helped me see your original valid point on harnessless LLMs! >You can achieve a lot by coding around the issues, but that's kinda against the point of 'AI', is it? Will think on that a bit more.