Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ddp26
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
61.
▲
by
ddp26
4mo ago
Author here. Agree, and I wrote in that section "Absolute accuracy is hard to compare across markets on one platform, and across platforms, because different forecasting questions have different difficulties. I addressed this by tracki
62.
▲
by
ddp26
5mo ago
Yeah, the question in the title can be answered: "by using gpt-4o, a model 2 years behind the frontier, to serve audio responses"
63.
▲
by
ddp26
5mo ago
Training window cutoff is Jan 2026, when Opus 4.6 was Aug 2025. That quite a lot of new world knowledge.
64.
▲
I think Anthropic is worth $100B more than last week
(futuresearch.ai)
9 points
by
ddp26
5mo ago
|
0 comments
65.
▲
by
ddp26
5mo ago
The free open source model does have its competitive advantages!
66.
▲
by
ddp26
5mo ago
The second paragraph starts "Muse Spark is the first step on our scaling ladder and the first product of a ground-up overhaul of our AI efforts. To support further scaling, we are making strategic investments..." This article is a
67.
▲
by
ddp26
6mo ago
Got a source on this? I didn't take into account in this forecast that public markets could be very inefficient in this way.
68.
▲
by
ddp26
6mo ago
There is actually a real bull case for xAI (that I don't endorse), e.g. from people who think that chips & computer is the main determiner of model quality. xAI may plausibly soon have the biggest training apparatus of anyone. I th
69.
▲
by
ddp26
6mo ago
As I wrote in the piece, I'm extremely skeptical that xAI should be valued as if it is a frontier lab. But as you say, going back to the xAI + SpaceX merger, analysts consistently seem to value it as if it is, so I predict the public w
70.
▲
by
ddp26
6mo ago
Yeah, I might have stated this poorly. In the forecast it's just a question of expected value, I don't give almost any probability to "Starship is worthless". My 50% CI on Starship's fair market value at IPO time is
71.
▲
by
ddp26
6mo ago
I read your comment as being glib, but in forecasting this I was really puzzled how much to anchor to how analysts tend to value these businesses. I ended up largely deferring to them, e.g. predicting the public will value xAI at $258 billi
72.
▲
by
ddp26
6mo ago
Yeah, it's wild. But it's not like the P/E should be 30, what do you think would be fair? That's the thing about SpaceX, some businesses are real businesses that can be modeled in normal ways, like the government launch
73.
▲
A forecast of the fair market value of SpaceX's businesses
(futuresearch.ai)
100 points
by
ddp26
6mo ago
|
205 comments
74.
▲
by
ddp26
6mo ago
Yeah, but uvx has this thing where it can automatically build the latest environment, and pull the latest (unpinned) version, right?
75.
▲
by
ddp26
6mo ago
My team was making fun of me for starting all my chats with "Hi Claude"
76.
▲
by
ddp26
6mo ago
Yeah, sharing information across Claude Code sessions really is a problem that needs solving. An urgent hack, where you're using Claude Code to debug and trying to get help from your team, is one such case.
77.
▲
by
ddp26
6mo ago
Yeah, and this is a pattern I saw in the Fancy Bear Goes Fishing book, a lot of discovery of malware is either pure luck, or blunders from the malware developers. https://en.wikipedia.org/wiki/Fancy_Bear_Goes_Phishing
78.
▲
by
ddp26
6mo ago
Agree, lots of hand wringing about us being so vulnerable to supply chain attacks, but this was handled pretty well all things considered
79.
▲
by
ddp26
6mo ago
Sure, but this is a pretty onerous restriction. Do you think supply chain attacks will just get worse? I'm thinking that defensive measures will get better rapidly (especially after this hack)
80.
▲
by
ddp26
6mo ago
Yeah, this was my team at FutureSearch that had the lucky experience of being first to hit this, before the malware was disclosed. One thing not in that writeup is that very little action was needed for my engineer to get pwnd. uvx automati
81.
▲
by
ddp26
6mo ago
I think so, and I've seen other solutions too. The one in the OP is more general, as you say. Have you tried Code Mode?
82.
▲
Ask LLM Agents to Classify Problems Before Starting
(futuresearch.ai)
7 points
by
ddp26
7mo ago
|
1 comments
83.
▲
by
ddp26
7mo ago
I don't understand the CLI vs MCP. In cli's like Claude Code, MCPs give a lot of additional functionality, such as status polling that is hard to get right with raw documentation on what APIs to call.
84.
▲
by
ddp26
7mo ago
Prediction markets are interesting when they are predicting future things nobody knows for sure. "Predicting" private, known information is the wrong use case.
85.
▲
by
ddp26
7mo ago
This has persisted for a crazy long time. It's one thing to ship your org chart, it's another thing to leave it in prod for years! Reminds me of the YouTube Music vs Google Play Music debacle.
86.
▲
by
ddp26
7mo ago
I tried using wolfram alpha as a tool for an llm research agent, and I couldn't find any tasks it could solve with it, that it couldn't solve with just Google and Python.
87.
▲
I ran 10,000 web research agents
(everyrow.io)
12 points
by
ddp26
7mo ago
|
0 comments
88.
▲
by
ddp26
7mo ago
How do you manage the laptop + mouse?
89.
▲
by
ddp26
8mo ago
Yep! We have lots of examples like that where two vendors, or two customers, are completely non-matching. With LLMs and LLM web agents, you also can associate things that are not the same entity. One example we have is merging a table of co
90.
▲
How LLM agents solve the table merging problem
(futuresearch.ai)
29 points
by
ddp26
8mo ago
|
3 comments
More ›