8 ms·
It was true 6 months ago, not anymore. Frontier models now outperform developers on many tasks, be it on quality/readability/maintainability, and let’s not talk
by noname120 4mo ago
It was true 6 months ago, not anymore. Frontier models now outperform developers on many tasks, be it on quality/readability/maintainability, and let’s not talk about speed…
- suddenlybananas 4mo agoWhy is anthropic hiring software developers then?
- aspenmartin 4mo agoBecause they still need them?
- suddenlybananas 4mo agoWhy would they still need them if "[f]rontier models now outperform developers on many tasks, be it on quality/readability/maintainability, and let’s not talk about speed"
- aspenmartin 4mo agoBecause to replace a SWE you need them to reliably outperform developers on ALL tasks
- suddenlybananas 4mo agoBut anthropic already had plenty of developers. Why would they actively need to hire more if the workload is all being automated?
- aspenmartin 4mo agoBecause it’s a force multiplier at this stage in the capability ladder. ROI of developers is arguably higher or will become higher in a matter of months.
- hajile 4mo agoThat's some really fast goalpost moving. If AI could outperform humans, Anthropic would NEVER release that model. Instead, they'd use it to create a new google, photoshop, office, windows, etc for cheap then undercut all those companies and taking over the entire software industry.
- lunar_mycroft 4mo agoWorth noting that the person you're replying to is not the same as the one who said frontier models now outperform developers.
- aspenmartin 4mo agoIt can outperform humans, just unevenly. You’ll see a lot of the same dynamics as you see with Mythos which, tbh is kind of refreshing. I get the sense that Dario while of course forced to ruthlessly run a company is genuinely interested in figuring out how to roll this out as ethically as he can.
- lunar_mycroft 4mo agoI've seen the code they produce without extensive help from human developers, this is clearly false. Good to see the classic "yeah the models weren't good enough six months ago, but this time they actually are, promise! Please forget you were hearing the exact same thing six months ago!" is alive and well though.
- aspenmartin 4mo agoAre you aware of performance trends though? You’re painting a picture that seems to ignore how things have consistently trended for many years now, even pre ChatGPT. It is absolutely data driven to say “an inflection point has happened within the last 6 months”. And that was also true 6 months ago (where people started using coding agents fairly consistently since sonnet 4). And it was true 6 months before that. It’s not like people are like “we’ve fixed all the bugs!” And then nothing has changed. I don’t necessarily agree with the parent poster that agents are better than humans but they are certainly much better at many tasks.
- lunar_mycroft 4mo ago> Are you aware of performance trends though? You’re painting a picture that seems to ignore how things have consistently trended for many years now, even pre ChatGPT. Models have been getting better, but all that follows from that is that newer models tend to be better than older ones. It doesn't follow that they have (or even will in the future) gotten better than anything else, be that human developers, a given definition of good enough, etc. > It is absolutely data driven to say “an inflection point has happened within the last 6 months”. With all due respect to OP (who I think is responsible for popularizing that way of phrasing it), I don't think it is when you consider the actual definition of "inflection point". At best I think you can say that models crossed a lot of developers definition of good enough around then, which is a different thing. The problem I have with that is that as a (mostly) outsider looking in, it doesn't seem like they're right.
- aspenmartin 4mo ago> Models have been getting better, but all that follows from that is that newer models tend to be better than older ones. It doesn't follow that they have (or even will in the future) gotten better than anything else, be that human developers, a given definition of good enough, etc. But this is not true, you’re saying we only have relative performance numbers and not absolute measures of capabilities and reliability but that’s simply not true. OSS benchmarks as well as the internal flywheels of these companies are good complementary measurements. > At best I think you can say that models crossed a lot of developers definition of good enough around then, which is a different thing That’s the inflection point. Implication is a massive jump in adoption. We’re not like pulling this out of a hat, there are a number of compelling datapoints. The onus is on people to bring actual evidence that contradicts all of the data and observations we have.