Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mdahardy
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
mdahardy
9mo ago
Our main argument is that outputs will become increasingly indistinguishable, but the processes won't. E.g. in 5 years if you watch an AI book a flight it will do it in a very non-human way, even if it gets the same flight you your
2.
▲
by
mdahardy
9mo ago
This is a fair criticism we should've addressed. There's actually a nice study on this: Vong et al. ( https://www.science.org/doi/10.1126/science.adi1374 ) hooked up a camera to a baby's head so it wo
3.
▲
AI capability isn't humanness
(research.roundtable.ai)
52 points
by
mdahardy
9mo ago
|
54 comments
4.
▲
by
mdahardy
9mo ago
Ah, nice idea. I hadn't considered locking it after you guess correctly.
5.
▲
Show HN: ModelGuessr: Can you tell which AI you're chatting with?
(model-guessr.com)
6 points
by
mdahardy
9mo ago
|
2 comments
6.
▲
by
mdahardy
10mo ago
The cross-tile challenges were quite robust - every model struggled with them, and we tried with several iterations of the prompt. I'm sure you could improve with specialized systems, but the models out-of-the-box definitely struggle w
7.
▲
by
mdahardy
10mo ago
That's a cool idea. I bet it would work better.
8.
▲
by
mdahardy
10mo ago
yes
9.
▲
by
mdahardy
10mo ago
After watching hundreds of these runs, Gemini was by far the least frustrating model to observe.
10.
▲
by
mdahardy
10mo ago
We have an example of a failed cross-tile result in the article - the models seem like they're much better at detecting whether something is in an image vs. identifying the boundaries of those items. This probably has to do with how th
11.
▲
by
mdahardy
10mo ago
Same! As we talk about in the article, the failures were less from raw model intelligence/ability than from challenges with timing and dynamic interfaces
12.
▲
by
mdahardy
10mo ago
While running this I looked at hundreds and hundreds of captchas. And I still get rejected on like 20% of them when I do them. I truly don't understand their algorithm lol
13.
▲
by
mdahardy
10mo ago
You could definitely do better than we do here - this was just a test of how well these general-purpose systems are out-of-the-box
14.
▲
Benchmarking leading AI agents against Google reCAPTCHA v2
(research.roundtable.ai)
124 points
by
mdahardy
10mo ago
|
97 comments
15.
▲
GPT-5 negotiates harder and better than Opus 4.1
(mdahardy.substack.com)
1 points
by
mdahardy
1y ago
|
0 comments
16.
▲
by
mdahardy
1y ago
Roundtable | https://roundtable.ai | On-site San Francisco, CA | Full-time Roundtable is a research and deployment company building the proof-of-human layer in digital identity. Roundtable seeks to research and build real-world
17.
▲
Benchmarking bot detection systems against modern AI agents
(research.roundtable.ai)
3 points
by
mdahardy
1y ago
|
0 comments
18.
▲
The Best Companies Don't Solve Problems
(substack.com)
2 points
by
mdahardy
1y ago
|
0 comments
19.
▲
by
mdahardy
1y ago
Co-founder of Roundtable here. I agree that better authentication methods for AI agents are needed. But right now bots and malicious agents are a real problem for anyone running sites with significant traffic. In the long run I don’t think
20.
▲
by
mdahardy
2y ago
Seems like a lot of this could be explained by better food tending to be served in locations with lower commercial real estate prices (I believe Tyler Cowen has written about this).
21.
▲
How we spot AI using keystrokes: Lessons from analyzing 5M+ survey responses
2 points
by
mdahardy
2y ago
|
0 comments
22.
▲
Forget No‐Code. The Future Is All‐Code (Thanks to LLMs)
(github.com)
5 points
by
mdahardy
2y ago
|
0 comments
23.
▲
by
mdahardy
2y ago
Thanks for letting us know - we're planning to support TypeScript soon. A lot of stuff on our roadmap!