5 ms·
Show HN: Benchmark your eng team's AI agent maturity in 5 minutes
we had hundreds of discussions with engineering leaders over the past few months, and everyone's trying to understand where they are in the AI journey.
we collected all this data into a benchmark and built a free grader to let you know where you stand.
you answer on a 1–5 scale (e.g., autonomy runs from "suggestions only" to "agents own multi-hour workflows across code, infra, and external systems") - takes about 5 minutes.
https://agent-benchmarks.com/software-factory/ https://agent-benchmarks.com/software-factory/
waiting for your results!
- bohaska 2mo ago[dead]
- dude250711 2mo agoIt's like sheep measuring which of them are closer to the front of a flock...
- adamgold7 2mo agowhat do you mean?
- nDRDY 2mo agoI got my agent to fill it out for me, but it won't tell me my score!
- xnx 2mo ago"software factory" seems like the wrong analogy. Factories make identical products over and over. Software is different every time (otherwise you just copy and paste). "Software machine shop" would be closer.
- adamgold7 2mo agothat's a very interesting take! i think the market is leaning towards software factory now to describe autonomous agents in R&D and that might very well not reprensent the truth
- t-writescode 2mo agoI got a high rating because I use AI-powered auto-complete, search and CodeRabbit for MRs. I'm suspicious. Perhaps the explanations provided by the hover-overs is less than clear?
- adamgold7 2mo agowhat were you expecting?
- t-writescode 2mo agoFar lower. I assumed most companies aren’t even doing that, especially with the options I was given. I was a 1 or a 2 for most things, and a few 4s as appropriate.
- zerolunier 2mo ago[dead]