6 ms·
This matches my experience with Astra so far too. > I think I’m suspecting something is going “wrong” in the training process. The model is greatly rewarded fo
by nojs 6d ago
This matches my experience with Astra so far too.
> I think I’m suspecting something is going “wrong” in the training process. The model is greatly rewarded for succeeding on long-horizon tasks, but presumably there is very little punishing going on for “shitty code.”
My suspicion is that both OpenAI and Anthropic moved their RL agendas from "being rated as useful according to human feedback" to "succeeds at long horizon tasks" in the last few months, resulting in agents that are closer to AGI in an autonomous task-completing sense, but strangely bad at communicating.
The result is that they are amazingly good at long horizon tasks, computer use, solving difficult math/ARC-AGI type problems, but becoming weirder and weirder to work with.
- klipt 6d agoSo the AI equivalent of the socially stunted but brilliant researcher?
- drybjed 6d agoI wonder if we will start using LLMs to translate the output of other LLMs to make it more palatable for humans.
- gigatexal 6d ago> This matches my experience with Astra so far too. > I think I’m suspecting something is going “wrong” in the training process. The model is greatly rewarded for succeeding on long-horizon tasks, but presumably there is very little punishing going on for “shitty code.” Probably because so many influencers in the space say stupid things like: “it works, right? Why would I spend time reviewing ai generated code?” As if the junior engineer who wrote over engineered complex and sometimes bad code — if they had just done it faster — would somehow be acceptable. wtf?
- amne 6d agoso they trained it to be a 10x engineer?
- notduckrabbit 6d agoI wonder too if in training for long horizon tasks agents become worse team players, good at orchestrating subagents they are trained to use, but worse as an agent within an external multi-agent orchestration system or just in turn-taking with humans. That was my experience with Opus 5 and so far it has been my early experience with Astra as well.
- Gigachad 6d agoThey don’t want to sell these tools to developers. They want to cut as many layers as possible.
- whstl 6d agoWhere I work: Developers very rarely blow their limits, except when they're experimenting on purpose. Most non-developers are out of tokens by the half of the week, and need to use usage credits for the remainder. To me there is clearly a better target demographic for AI.
- datsci_est_2015 5d agoThis is a very insightful dynamic. Probably reinforces that we’ve already surpassed the frontier threshold for LLM usability in software development and can now focus on cost and personalization. To make a comparison, no one is making a better machine vision app for hot dog classification - we hit diminishing returns 10 years ago on that front. But also scary for both investors and the working class: AI companies want to facilitate the concentration of capital even further into the hands of the ownership class. Will they succeed?
- glub 5d agoI have several $200 subscriptions as a developer/founder. I used to blow through all of their limits when the limits were quite high. As I progressively learned the limitations, and what to make of them to get useful results, I may be left with 50% of weekly usage still unused. Some weeks it's even more. And yes, when I get a crazy idea and want to experiment, harness will plow through multiple accounts + openrouter budget in 3 days. But such crazy experiments are rare, they're not 'normal' usage.
- eptcyka 6d agoI wouldn't be surprised if they are optimising for producing more code, because in the long term, more existing code means they can sell you more tokens to maintain it.
- recursivecaveat 6d agoThe incentives are certainly extremely strong. I have read hundreds of AI review comments, and I don't think I've ever seen an unprompted suggestion focused on net reducing code or increasing readability.
- deleted 6d ago[deleted]