Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
janaksunil
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
janaksunil
4d ago
do you have some time to chat? janak@withspecific.com
2.
▲
by
janaksunil
4d ago
would love to chat and learn more about your set up! here's my email - janak@withspecific.com
3.
▲
by
janaksunil
4d ago
we manually vet all codebases and companies
4.
▲
by
janaksunil
4d ago
yes!
5.
▲
by
janaksunil
4d ago
that's fair - for long horizon engineering tasks would speed still matter?
6.
▲
by
janaksunil
4d ago
this is a good question. what would make you reject an otherwise working PR on design grounds?
7.
▲
by
janaksunil
4d ago
we reached out to companies that were willing to license their codebases. every codebase we used had real users, one of them had 200k+ users and is currently top 100 on the app store.
8.
▲
by
janaksunil
4d ago
i appreciate the feedback, the benchmark is primarily long horizon real world engineering tasks on big private codebases.
9.
▲
by
janaksunil
4d ago
we're going to open source some of our tasks and model trajectories as well
10.
▲
by
janaksunil
4d ago
all the codebases were written pre-2023, so pre when AI got good at coding
11.
▲
by
janaksunil
4d ago
would love to learn why?
12.
▲
by
janaksunil
4d ago
will do! happy to chat more on janak@withspecific.com as well
13.
▲
by
janaksunil
4d ago
we've done our best to use the native provider's harness. all models were run on 'high' reasoning. this is still v1 and tons of room for improvement - really appreciate your feedback!
14.
▲
by
janaksunil
4d ago
makes sense
15.
▲
by
janaksunil
4d ago
the tasks on the benchmar are long horizon swe tasks - where GLM does surprisingly well
16.
▲
by
janaksunil
4d ago
here's where all the models messed up! its under this section 'Missed requirements are the most common failure' on realswe.withspecific.com we also have the setup in the blog. the reason for lower success rates is that we gav
17.
▲
by
janaksunil
4d ago
nope, these were private codebases
18.
▲
by
janaksunil
4d ago
hey this seems really interesting - what prompted you to test multiple agents on your codebase?
19.
▲
by
janaksunil
4d ago
i'm janak, cofounder of Specific Labs (YC F25) and one of the authors of Real-SWE. if i can help answer any questions please feel free to email me at janak@withspecific.com, happy to send over my phone number as well :)
20.
▲
Show HN: MCP server that finds dev tool credits in your workflow
1 points
by
janaksunil
6mo ago
|
1 comments