Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
marsh_mellow
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
NanoClaw solves one of OpenClaw's biggest security issues
(venturebeat.com)
46 points
by
marsh_mellow
7mo ago
|
28 comments
2.
▲
Benchmarking GPT-5 on 400 real-world code reviews
(qodo.ai)
72 points
by
marsh_mellow
1y ago
|
80 comments
3.
▲
Kubernetes Will Solve YAML Headaches with Kyaml
(thenewstack.io)
4 points
by
marsh_mellow
1y ago
|
0 comments
4.
▲
Matter vs. Force: Why There Are Two Types of Particles
(quantamagazine.org)
2 points
by
marsh_mellow
1y ago
|
0 comments
5.
▲
Improving Deep Learning with a Little Help from Physics
(quantamagazine.org)
10 points
by
marsh_mellow
1y ago
|
0 comments
6.
▲
by
marsh_mellow
1y ago
p-value of 7.9% — so very close to statistical significance. the p-value for GPT-4.1 having a win rate of at least 49% is 4.92%, so we can say conclusively that GPT-4.1 is at least (essentially) evenly matched with Claude Sonnet 3.7, if not
7.
▲
by
marsh_mellow
1y ago
I don't think the absolute score means much — judge models have a tendency to score around 7/10 lol 55% vs. 45% equates to about a 36 point difference in ELO. in chess that would be two players in the same league but one with a cl
8.
▲
by
marsh_mellow
1y ago
Good point. They said they validated the results by testing with other models (including Claude), as well as with manual sanity checks. 55% to 45% definitely isn't a blowout but it is meaningful — in terms of ELO it equates to about a
9.
▲
by
marsh_mellow
1y ago
From OpenAI's announcement: > Qodo tested GPT‑4.1 head-to-head against Claude Sonnet 3.7 on generating high-quality code reviews from GitHub pull requests. Across 200 real-world pull requests with the same prompts and conditions, th
10.
▲
Determining IaC ownership – a tag-based approach
(token.security)
5 points
by
marsh_mellow
1y ago
|
6 comments
11.
▲
Expanded language support in Amazon Q Developer
(aws.amazon.com)
1 points
by
marsh_mellow
1y ago
|
0 comments
12.
▲
by
marsh_mellow
2y ago
Very cool! Could this work for detecting nearby drones?
13.
▲
Computer Use API Documentation
(docs.anthropic.com)
13 points
by
marsh_mellow
2y ago
|
2 comments
14.
▲
by
marsh_mellow
2y ago
Anthropic blog post outlining the research process: https://www.anthropic.com/news/developing-computer-use Computer use API documentation: https://docs.anthropic.com/en/docs/build-with-claude&
15.
▲
The U.S. AI Safety Institute stands on shaky ground
(techcrunch.com)
2 points
by
marsh_mellow
2y ago
|
0 comments
16.
▲
Developing a Computer Use Model
(anthropic.com)
3 points
by
marsh_mellow
2y ago
|
1 comments
17.
▲
AI Agents: A Comprehensive Introduction for Developers
(thenewstack.io)
1 points
by
marsh_mellow
2y ago
|
0 comments
18.
▲
How We Generated Millions of Content Annotations
(engineering.atspotify.com)
3 points
by
marsh_mellow
2y ago
|
0 comments
19.
▲
Open-source tool for LLM document processing pipelines from UC Berkeley
(github.com)
1 points
by
marsh_mellow
2y ago
|
0 comments
20.
▲
by
marsh_mellow
2y ago
To tag on to this, what are the most useful capabilities besides code generation?
21.
▲
by
marsh_mellow
2y ago
This is great. Is there any work being done to make something similar part of the browser API?
22.
▲
by
marsh_mellow
2y ago
There's an open source version of this as well: https://github.com/Codium-ai/pr-agent
23.
▲
PR-Agent — extension that adds AI chat to code reviews on GitHub
(chromewebstore.google.com)
24 points
by
marsh_mellow
2y ago
|
15 comments
24.
▲
by
marsh_mellow
2y ago
They list seven different use cases in this technical blog: https://harrison.ai/news/reimagining-medical-ai-with-the-mos... I'd interpret it as a foundation model in the radiology domain
25.
▲
US Government Sets Out to Improve Internet Routing Security
(infosecurity-magazine.com)
41 points
by
marsh_mellow
2y ago
|
8 comments
26.
▲
Pinot for Low-Latency Offline Table Analytics
(uber.com)
1 points
by
marsh_mellow
2y ago
|
0 comments