Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ddp26
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
ddp26
3mo ago
If it's helpful, I'm still holding at July 9 as my median date that Fable gets re-released to Americans, the news of the last 24 hours didn't update the model meaningfully.
32.
▲
by
ddp26
3mo ago
People forget that Meta already did this years ago, before prediction markets became the next big consumer trend for them to chase. The app was called Forecast, and launched in June 2020. (Around the same time that Kalshi and Polymarket lau
33.
▲
World-Modeling the US vs. Anthropic on Claude Fable
(lesswrong.com)
9 points
by
ddp26
3mo ago
|
1 comments
34.
▲
by
ddp26
3mo ago
You mean chatgpt style AI won't help them with those skills? If a human parent or teacher can help with skills like reading, an AI system can too, once it's trained and designed to do so. (How good are humans at teaching reading a
35.
▲
by
ddp26
3mo ago
Yes, I have, but comments are still useful. Don't you think this is an overcorrection?
36.
▲
Conscripting engineers to make training data won't push AI
(futuresearch.ai)
1 points
by
ddp26
3mo ago
|
0 comments
37.
▲
How the US vs. Anthropic Standoff on Claude Fable Will End
(futuresearch.ai)
2 points
by
ddp26
3mo ago
|
1 comments
38.
▲
Claude can miss the motives of politicians
(futuresearch.ai)
10 points
by
ddp26
3mo ago
|
0 comments
39.
▲
Measuring one way AIs lack self-awareness
(futuresearch.ai)
1 points
by
ddp26
4mo ago
|
0 comments
40.
▲
by
ddp26
4mo ago
It is refreshing but perhaps actually not warranted this time? I mostly study web research, and Opus 4.7 was a regression on BrowseComp compared to Opus 4.6, which has been born out by my usage. Opus 4.8 is now much better than either 4.7 o
41.
▲
by
ddp26
4mo ago
What's a definition of AGI you would use, for either time, tasks, value, or job descriptions?
42.
▲
by
ddp26
4mo ago
I linked elsewhere in a comment, Metaculus has AGI forecasts. You can also now use AI forecasters like FutureSearch [1] (disclaimer: I work there), which are competitive with the best humans / teams of humans. And since you aren't
43.
▲
by
ddp26
4mo ago
Thank you! Tok me a few hours, without Claude Code I don't think I would have even attempted this.
44.
▲
by
ddp26
4mo ago
It's been a big problem for a while. The big Metaculus question about AGI has depends on the game "Montezuma's revenge" (!), and there have been many debates about this going back to at least 2020: https://www
45.
▲
by
ddp26
4mo ago
Author here, I agree, I'd be happy if admins want to change the title of this submission to the title of the piece.
46.
▲
by
ddp26
4mo ago
Author here, I drew on this from AI 2027. Yes, a very-expensive AGI, e.g. $1 million / day to simulate a smart human, would be a huge deal. But it would have meaningfully different effects than a cheap one. Here's one definition A
47.
▲
by
ddp26
4mo ago
I see a lot of comments like this is the blocking of prediction markets about politics, war, etc. It's important to remember that ~80% of activity Polymarket and ~90% of Kalshi, by volume, are sports. These are effectively sports betti
48.
▲
Some rare examples of AIs being underconfident
(futuresearch.ai)
6 points
by
ddp26
4mo ago
|
0 comments
49.
▲
by
ddp26
4mo ago
Snake oil is a bit strong, no? I would agree that the burden of proof is on multi-agent systems to show they are outperforming single-agent systems. On my own evals I have seen this, though the improvement may not have been worth the extra
50.
▲
by
ddp26
4mo ago
I like this, though it does leave me feeling more nervous when I really don't know how I'd solve the problem, still requires trust.
51.
▲
History doesn't repeat itself as often as LLMs think
(futuresearch.ai)
1 points
by
ddp26
4mo ago
|
0 comments
52.
▲
by
ddp26
4mo ago
It would have to be an incredibly tiny tax, no?
53.
▲
by
ddp26
4mo ago
What was your use case?
54.
▲
Agents Sometimes Catastrophize
(futuresearch.ai)
9 points
by
ddp26
4mo ago
|
2 comments
55.
▲
by
ddp26
4mo ago
Stanford has this policy too. Students get livid when proctoring is proposed, even though cheating is rampant (afaict)
56.
▲
Run Agents Twice
(futuresearch.ai)
6 points
by
ddp26
4mo ago
|
0 comments
57.
▲
by
ddp26
4mo ago
Author here. Great point, and I think this is due to what another commenter points out, that the questions are different. The right test of this is to take the _same_ markets that run for 90+ days, and check accuracy 90 days out vs 30 days
58.
▲
by
ddp26
4mo ago
Author here. Hal Varian pointed me to this 1992 paper, which I think is still considered the canonical empirical piece on what is actually going on in trading behavior that leads to accuracy (or not): https://www.jstor.org/s
59.
▲
by
ddp26
4mo ago
Yeah. People have put together a Prediction Market Database [1] (in a Google sheet), I think it's pretty well sourced and shows a good number of both real money and play money prediction markets from before 2002. DARPA did have a big r
60.
▲
by
ddp26
4mo ago
It's true they are "just" summarizing current knowledge. But there are better and worse summaries of current knowledge! Some summaries, like on some prediction markets, have objective accuracy that is much better than chance.
More ›