Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zone411
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
zone411
5d ago
Completely misleading. This is all you need to read and understand for Anthropic's FLT formalization: import Mathlib import Theorems.Thm_fermat_last_theorem /-- Solution side: the same statement, binder for binder, proved
2.
▲
by
zone411
13d ago
I don't: "Problems solved before a model's training cutoff can be filtered out, and all models compared on the remaining problems" means that the problems an older model actually solved are the ones that get filtered out
3.
▲
by
zone411
1mo ago
A complete mischaracterization, as usual for HN lately when discussing AI or LessWrong. Obviously, even average levels of persuasion are enough to convince some people. And nobody is air-gapping AI.
4.
▲
by
zone411
2mo ago
So why don't companies in other industries rush to prove their products are dangerous weapons? Maybe because it would be a really dumb PR stunt?
5.
▲
by
zone411
2mo ago
And how do people saying this know the capabilities of yet unreleased models?
6.
▲
by
zone411
2mo ago
How is it in their interest? Scaring customers, worrying employees, and inviting regulators to act is in their interest?
7.
▲
by
zone411
2mo ago
It's not good marketing for them. This is a talking point with zero evidence that people repeat mindlessly. Scaring customers, worrying employees, and inviting regulators to act would be the worst marketing idea ever devised.
8.
▲
Natural-Density Almost-Bounded Collatz Orbits in Logarithmic Time (AI, Lean)
(proofatlas.ai)
2 points
by
zone411
2mo ago
|
0 comments
9.
▲
by
zone411
3mo ago
Yes, definitely not a new idea. I had a multi-turn composite model in 2024 that was outperforming the top models across benchmarks: https://x.com/LechMazur/status/1828804485033992514 .
10.
▲
by
zone411
3mo ago
That's not proof. Emergent intelligence is not consciousness.
11.
▲
by
zone411
4mo ago
I’ve tested this model on four of my benchmarks: https://github.com/lechmazur/buyout_game 10th out 36. https://github.com/lechmazur/pact/ 14th out 25. https://github.com/lechm
12.
▲
by
zone411
4mo ago
100%. It's sad to see that this attitude has spread to HN
13.
▲
by
zone411
4mo ago
I actually tried using GPT-5.5 Pro on this problem recently. It thought it was making progress on one path, but it made so many mistakes that it didn't feel worth it pushing further. It'll be interesting to check whether it's
14.
▲
LLM Position Bias Benchmark: Swapped-Order Pairwise Judging
(github.com)
1 points
by
zone411
5mo ago
|
0 comments
15.
▲
by
zone411
5mo ago
https://variety.com/2020/digital/news/twitter-unblocks-new-y...
16.
▲
Show HN: Buyout Game Benchmark: Multi-Agent Bargaining, Transfers, and Takeovers
(github.com)
6 points
by
zone411
6mo ago
|
0 comments
17.
▲
by
zone411
6mo ago
I built this benchmark this month: https://github.com/lechmazur/sycophancy . There are large differences between LLMs. There are large differences between LLMs. For example, Mistral Large 3 and GPT-4.1 will initially ag
18.
▲
by
zone411
6mo ago
I built two related benchmarks this month: https://github.com/lechmazur/sycophancy and https://github.com/lechmazur/persuasion . There are large differences between LLMs. For example, good luck get
19.
▲
LLM Persuasion Benchmark: Multi-Turn Persuasion Between Models
(github.com)
9 points
by
zone411
6mo ago
|
0 comments
20.
▲
by
zone411
6mo ago
Hmm, maybe in the next edition, Opus gets expensive. I should probably run GPT-5.4 xhigh too if I do that for fairness...
21.
▲
Show HN: LLM Debate Benchmark
(github.com)
9 points
by
zone411
6mo ago
|
3 comments
22.
▲
by
zone411
6mo ago
Rationalists were right about everything that mattered: crypto, AI, COVID... HN commentators, by contrast, were wrong about everything that mattered.
23.
▲
Show HN: LLM Sycophancy Benchmark: Opposite-Narrator Contradictions
(github.com)
3 points
by
zone411
6mo ago
|
0 comments
24.
▲
by
zone411
7mo ago
Results from my Extended NYT Connections benchmark: GPT-5.4 extra high scores 94.0 (GPT-5.2 extra high scored 88.6). GPT-5.4 medium scores 92.0 (GPT-5.2 medium scored 71.4). GPT-5.4 no reasoning scores 32.8 (GPT-5.2 no reasoning scored 28.1
25.
▲
by
zone411
7mo ago
I've made top-10 lists of LLMs' favorite names to use in creative writing here: https://x.com/LechMazur/status/2020206185190945178 . They often recur across different LLMs. For example, they love Elara an
26.
▲
by
zone411
7mo ago
They're improved compared to 4.5 on my Extended NYT Connections benchmark ( https://github.com/lechmazur/nyt-connections/ ). Sonnet 4.6 Thinking 16K scores 57.6 on the Extended NYT Connections Benchmark. Sonnet
27.
▲
by
zone411
8mo ago
For people interested in these kinds of benchmarks, I have two multiplayer, multi-round games: - Elimination Game Benchmark: Social Reasoning, Strategy, and Deception in Multi-Agent LLM Dynamics at https://github.com/lechmaz
28.
▲
by
zone411
9mo ago
Scores 92.0 on my Extended NYT Connections benchmark ( https://github.com/lechmazur/nyt-connections/ ). Gemini 2.5 Flash scored 25.2, and Gemini 3 Pro scored 96.8.
29.
▲
by
zone411
9mo ago
I've benchmarked it on the Extended NYT Connections benchmark ( https://github.com/lechmazur/nyt-connections/ ): The high-reasoning version of GPT-5.2 improves on GPT-5.1: 69.9 → 77.9. The medium-reasoning vers
30.
▲
by
zone411
10mo ago
I haven't looked in the logs for this in this particular project, but I've seen this occur frequently in my multiplayer benchmarks.
More ›