Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
narush
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
narush
25d ago
Cool post! I'm a co-author on a recent blog post from METR about the NanoGPT speed-run here [1]. I think it'd be of interest to anyone who enjoyed the original post. Appreciate the good beefy runs and spend here, it's a (from
2.
▲
by
narush
1y ago
We call this over-generalization out specifically in the "We do not provide evidence that:" table in the blog post and paper - I agree there are tasks these developers are likely sped up on with early-2025 tools.
3.
▲
by
narush
1y ago
Hey, thanks for linking this! I'm a study author, and I greatly appreciate that this author dug into the appendix and provided feedback so that other folks can read it as well. A few notes if it's helpful: 1. This post is primaril
4.
▲
by
narush
1y ago
Hey, thanks for digging into the details here! Copying a relevant comment ( https://news.ycombinator.com/item?id=44523638 ) from the other thread on the paper, in case it's help on this point. 1. Some prior studies that
5.
▲
by
narush
1y ago
Thanks for the feedback! I strongly agree this is not the only measure of developer productivity -- but it's certainly one of them. I think this measure as speaks very directly to how _many_ developers (myself included) understand the
6.
▲
by
narush
1y ago
Hey HN -- study author here! (See previous thread on the paper here [1].) I think this blog post is an interesting take on one specific factor that is likely contributing to slowdown. We discuss this in the paper [2] in the section "Im
7.
▲
by
narush
1y ago
Check out section AI increasing issue scope (C.2.3) in the paper -- https://metr.org/Early_2025_AI_Experienced_OS_Devs_Study.pdf We speak (the best we can) to changes in amount of code -- I'll note that this metric is
8.
▲
by
narush
1y ago
How these results transfer to other settings is an excellent question. Previous literature would suggest speedup -- but I'd be excited to run a very similar methodology in those settings. It's already challenging as models + tools
9.
▲
by
narush
1y ago
Sorry, this is the first 8 issues per-developer!
10.
▲
by
narush
1y ago
Thank you!
11.
▲
by
narush
1y ago
We attempted to! We explore this more in the section Trading speed for ease (C.2.5) in the paper ( https://metr.org/Early_2025_AI_Experienced_OS_Devs_Study.pdf ). TLDR: mixed evidence that developers make it less effortful, f
12.
▲
by
narush
1y ago
There's additional breakdown per-minute in the appendix -- see appendix E.4!
13.
▲
by
narush
1y ago
You can see a list of repositories with participating developers in the appendix! Section G.7. Paper is here: https://metr.org/Early_2025_AI_Experienced_OS_Devs_Study.pdf
14.
▲
by
narush
1y ago
You can see this analysis in the factor analysis of "Below-average use of AI tools" (C.2.7) in the paper [1], which we mark as an unclear effect. TLDR: over the first 8 issues, developers do not appear to get majorly less slowed d
15.
▲
by
narush
1y ago
Thanks for the kind words!
16.
▲
by
narush
1y ago
The instructions given to developers was not just "implement with AI" - but rather that they could use AI if they deemed it would be helpful, but indeed did _not need to use AI if they didn't think it would be helpful_. In ab
17.
▲
by
narush
1y ago
Honestly, this is a fair point -- and speaks the difficulty of figuring out the right baseline to measure against here! If we studied folks with _no_ AI experience, then we might underestimate speedup, as these folks are learning tools (see
18.
▲
by
narush
1y ago
Yep, sorry, meant to post this somewhere but forgot in final-paper-polishing-sprint yesterday! We'll be releasing anonymized data and some basic analysis code to replicate core results within the next few weeks (probably next, dependin
19.
▲
by
narush
1y ago
> which feels like it is easier and hence faster. We explore this factor in section (C.2.5) - "Trading speed for ease" - in the paper [1]. It's labeled as a factor with an unclear effect, some developers seem to think so,
20.
▲
by
narush
1y ago
The graphs are all matplotlib. The methodology figure is built in Figma! (Source: I'm a paper author :)).
21.
▲
by
narush
1y ago
Yeah, I'll note that this study does _not_ capture the entire OS dev workflow -- you're totally right that reviewing PRs is a big portion of the time that many maintainers spend on their projects (and thanks to them for doing this
22.
▲
by
narush
1y ago
Qualitatively, we don't see a drop in PR quality in between AI-allowed and AI-disallowed conditions in the study; the devs who participate are generally excellent, know their repositories standards super well, and aren't really in
23.
▲
by
narush
1y ago
Sounds great. Looking forward to hearing more detailed thoughts -- my emails in the paper :)
24.
▲
by
narush
1y ago
Noting that most of our power comes from the number of tasks that developers complete; it's 246 total completed issues in the course of this study -- developers do about 15 issues (7.5 with AI and 7.5 without AI) on average.
25.
▲
by
narush
1y ago
Hey Simon -- thanks for the detailed read of the paper - I'm a big fan of your OS projects! Noting a few important points here: 1. Some prior studies that find speedup do so with developers that have similar (or less!) experience with
26.
▲
by
narush
1y ago
Our largest funding was through The Audacious Project -- you can see an announcement here: https://metr.org/blog/2024-10-09-new-support-through-the-aud... Per our website, “To date, April 2025, we have not accepted com
27.
▲
by
narush
1y ago
Hey HN, study author here. I'm a long-time HN user -- and I'll be in the comments today to answer questions/comments when possible! If you're short on time, I'd recommend just reading the linked blogpost or the anno
28.
▲
by
narush
2y ago
I’ve replicated the OthelloGPT results mentioned in this paper personally - and it def felt like the next-move-only accuracy metric was not everything. Indeed, the authors of the original paper knew this, and so further validated the world
29.
▲
by
narush
2y ago
Hey Jeff, thanks for the kind words. I'd love to learn more about your experience and transition through that pain point -- shoot me an email at nate @ sagacollab . com if you want to chat.
30.
▲
by
narush
2y ago
> To be clear, though, you don't necessarily have parity of outputs. The cool thing is that the Excel file is both the programatic specification of the process as well as the actual output data you want as well. We can check parity
More ›