Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
usaar333
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
usaar333
13d ago
It's sub-Fable 5 on mirror code: https://epoch.ai/benchmarks/mirrorcode?view=graph&tab=leader...
2.
▲
by
usaar333
24d ago
It was in one paragraph mentioned, without much detail. But yes, I agree it is one of the largest barriers.
3.
▲
by
usaar333
1mo ago
Where's that stat from? https://www.pewresearch.org/internet/2026/06/17/americans-an... That's 30% positive on productive/informed, and only ~5% on hinders. Whether you think AI will turn
4.
▲
by
usaar333
2mo ago
> and eventually AI model will be commodified This axiom not being true (and I'd bet against it) means your overall conclusion is false.
5.
▲
by
usaar333
4mo ago
Tech workers get paid in equity and many in the semiconductor industry are making far far more than this a year with all the equity appreciation.
6.
▲
by
usaar333
4mo ago
How does delaying the release not solve anything? It puts everyone on a notice to fix all security vulnerabilities now
7.
▲
by
usaar333
5mo ago
I hate waiting on hold for 30 minutes even more.
8.
▲
by
usaar333
5mo ago
There's literally a link on the blog post to an article noting they hit $150M ARR.
9.
▲
by
usaar333
5mo ago
Voice agents have capabilities and policy to alter customer state. Just the other day I called into a CC company and the AI waived an interest charge.
10.
▲
by
usaar333
5mo ago
page is updated to state: MCP-Atlas: The Opus 4.6 score has been updated to reflect revised grading methodology from Scale AI.
11.
▲
by
usaar333
5mo ago
> But even setting aside the leaked answers, the scorer’s normalize_str function strips ALL whitespace, ALL punctuation, and lowercases everything before comparison. This means: I don't understand the concern here
12.
▲
by
usaar333
7mo ago
True, but it gets you higher accuracy. Gemini had the best aa-omniscience score https://artificialanalysis.ai/evaluations/omniscience
13.
▲
by
usaar333
7mo ago
Openai has; they don't even mention score on gpt-5.3-codex. On the other hand, it is their own verified benchmark, which is telling.
14.
▲
by
usaar333
7mo ago
i'd interpret that as rounding error. that is unchanged swe-bench seems really hard once you are above 80%
15.
▲
by
usaar333
10mo ago
In Quebec it was a 20% jump in mother employment: https://www.bloomberg.com/news/articles/2018-12-31/affordabl... And had all sorts of negative outcomes for the kids: https://www.edweek.org/te
16.
▲
by
usaar333
10mo ago
claude 4.5 gets 82% on their own highly customized scaffolding. (parallel compute with a scoring function). That beats Doubao
17.
▲
by
usaar333
1y ago
That wasn't a ceasefire violation. It was a six week ceasefire that had expired at the beginning of March
18.
▲
by
usaar333
1y ago
Physics seems better than veo 3 at least from demo videos
19.
▲
by
usaar333
1y ago
Except it is sublinear. Sonnet 4 was 10.2% above sonnet 3.7 after 3 months.
20.
▲
by
usaar333
1y ago
No it doesn't. If it were even linear compared to o1 -> o3, we'd be at 2.43 hours. Instead we're only at 2.29. Exponential would be at 3.6 hours
21.
▲
by
usaar333
1y ago
No, this is below expectations on both Manifold and lesswrong ( https://www.lesswrong.com/posts/FG54euEAesRkSZuJN/ryan_green... ). Median was ~2.75 hours on both (which already represented a bearish slowdown). Not
22.
▲
by
usaar333
1y ago
At this point the prediction for SWE bench (85% by end of this month) is not materializing. We're actually quite far away.
23.
▲
by
usaar333
1y ago
No obvious gains I feel from quick chats, but too early to tell. These benchmark gains aren't that high, so I doubt it is that obvious.
24.
▲
by
usaar333
1y ago
> Firstly, if your prior is that every previous startup failed, what does that say about your future chances of success? The prior is the market. It isn't sane to use your own prior experience. (Works both ways -- if your last start
25.
▲
by
usaar333
1y ago
Why is modal return so important? You'll work more than 2 jobs
26.
▲
by
usaar333
1y ago
It's a probabilistic model. It assumes (correctly) that the low probability of a home run times the home run's valuation is quite large ("expected returns" in the probabilistic sense). > this argument reads to me li
27.
▲
by
usaar333
1y ago
The value of the equity package is 4x higher than the FAANG equivalent equity package (at preferred/market pricing) - that's not the same as saying the shares themselves are worth that. To sum up the arguments: * Employment packag
28.
▲
Startup equity is worth more than you think
(amafinance.org)
2 points
by
usaar333
1y ago
|
9 comments
29.
▲
by
usaar333
1y ago
I don't see why the market cap proves whether she is correct or not. You'd have to compare it to the counter-factual of what the value of a Figma subsidiary would be under Adobe today. This is not obvious at all to me. Instagram
30.
▲
by
usaar333
1y ago
$19.8 billion market cap to save everyone from doing research
More ›