Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
brammertottens
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
24 ms
·
1.
▲
by
brammertottens
3mo ago
This is super interesting, and I like the idea of verifiable artifacts that an agent can produce, i.e. notebooks for analysis, links to the source for some claims. Building for scale, it would be interesting to know how the author thinks ab
2.
▲
by
brammertottens
3mo ago
I mean, they basically did that to make sure that Composer 2.5 sits on the right side, and to a quick reader looks good
3.
▲
by
brammertottens
3mo ago
This is an interesting finding, but very specialised. It would also be great to get some more information about the benchmark. Is it just a collection of files with vulnerabilities, or are they hidden in a real codebase, where LLM based app
4.
▲
by
brammertottens
3mo ago
It's an interesting post, but i'm a bit skeptical on their decision to report the best run for each agent, and not just the mean over the 5 runs. We have seen this as well in running benchmarks, that variance within one setup can
5.
▲
by
brammertottens
3mo ago
Just a question on the benchmark. It states that it is on real world code, but all the repos in the dataset are intentionally vulnerable repos right, not real world codebases that have reported vulnerabilities?
6.
▲
by
brammertottens
3mo ago
It goes a bit up and down, compared to two years ago, in my feeling, both anthropic and open ai coding models have made massive jumps. In between big releases I do feel the quality of the models varies over time. 2 years ago I got annoyed w
7.
▲
by
brammertottens
4mo ago
There definitely is the danger of a lot of garbage being shipped, but with both the models getting better and better, and more tools and ways of working being discovered, I believe the quality of what is being outputted is going up as well.