Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mpavlov
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
A Workaphile's Apology
(moalquraishi.wordpress.com)
3 points
by
mpavlov
1mo ago
|
0 comments
2.
▲
by
mpavlov
2mo ago
You mean all the possible states of the each element of the interface (buttons, forms, content blocks, etc.)?
3.
▲
by
mpavlov
2mo ago
Unfortunately, mainly because there're only self-reported and partial numbers for the benchmarks that matter the most.
4.
▲
by
mpavlov
2mo ago
> If we really want to benchmark the ability of models to use human UIs to solve problems, then perhaps we need to choose benchmarks that don’t have APIs available such that the model cannot get “creative” in any way and must use the UI
5.
▲
by
mpavlov
2mo ago
There's an anecdotal paper 'How We Broke Top AI Agent Benchmarks: And What Comes Next' https://moogician.github.io/blog/2026/trustworthy-benchmarks...
6.
▲
by
mpavlov
2mo ago
Haha, one day, one day...
7.
▲
You can't solve computer use by ignoring the interface
(steelmanlabs.com)
75 points
by
mpavlov
2mo ago
|
35 comments
8.
▲
Why Alpha Arena was a bad benchmark
(borisagain.substack.com)
6 points
by
mpavlov
11mo ago
|
0 comments
9.
▲
by
mpavlov
11mo ago
(author of PokerBattle is here) Well, you're not wrong :) Vercel is not the one to blame here, it's my skill issue. Entire thing was vibecoded by me — product manager with no production dev experience. Not to promote vibecoding, b
10.
▲
by
mpavlov
11mo ago
(author of PokerBattle here) You right, results and numbers are mainly for entertainment purposes. This sample size would allow to analyze main reasoning failure modes and how often they occur.
11.
▲
by
mpavlov
11mo ago
(author of PokerBattle here) Haven't seen it before, thanks Are you affiliated with them?
12.
▲
by
mpavlov
11mo ago
(author of PokerBattle here) I think it would've completely crush them (like any other solver-based solution). Poker is safe for now :)
13.
▲
by
mpavlov
11mo ago
(author of PokerBattle here) I noticed the same and think that you're absolutely right. I've thought about adding their current hand / draw, but it was too close to the event to test it properly.
14.
▲
by
mpavlov
11mo ago
(author of PokerBattle here) That’s true. The original goal was to see which model performs statistically better than the others, but I quickly realized that would be neither practical nor particularly entertaining. A proper benchmark would
15.
▲
by
mpavlov
11mo ago
(author of PokerBattle here) That's cool! Do you have a recording of the talk? You can use PokerKit ( https://pokerkit.readthedocs.io/en/stable/ ) for the engine.
16.
▲
by
mpavlov
11mo ago
(author of the PokerBattle here) Depends on what your goal is, I think. And it's also a thing — https://huskybench.com/
17.
▲
Show HN: Pokerbattle.ai – A week-long poker tournament for LLMs
(pokerbattle.ai)
14 points
by
mpavlov
1y ago
|
4 comments