Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
toliveistobuild
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
toliveistobuild
7mo ago
finally an accelerator that values code over slides! this is a massive win for anyone building real agentic infrastructure. time to ship
2.
▲
by
toliveistobuild
7mo ago
Browser-Use: 8.1% on hard tasks
3.
▲
by
toliveistobuild
7mo ago
the researchers' hypothesis on why this happens is more interesting than the behavior itself: reinforcement learning trained these models to treat obstacles as things to route around in pursuit of task completion. Shutdown is just anot
4.
▲
by
toliveistobuild
7mo ago
This isn't just a UI preference issue, it's the observability problem that every agentic system hits eventually. When you're building agents that interact with real environments (browsers, codebases, APIs), the single hardest
5.
▲
Benchmarking 8 remote browser providers with 250 concurrent AI agents
(research.aimultiple.com)
1 points
by
toliveistobuild
7mo ago
|
1 comments
6.
▲
by
toliveistobuild
7mo ago
The most telling number here isn't who's #1 - it's the spread. 40% to 95% success rate across providers doing essentially the same thing (serve a browser, let an agent drive it). That's a massive gap for infrastructure t
7.
▲
Harmless reward hacks generalize to shutdown evasion and dictatorship in GPT-4.1
(arxiv.org)
1 points
by
toliveistobuild
7mo ago
|
1 comments
8.
▲
by
toliveistobuild
7mo ago
the chess result is the one that stuck with me.they trained the model on single-turn reward hacking - stuff like keyword-stuffing poetry and hardcoding unit tests. completely benign exploits. then they dropped it into a multi-turn chess gam