Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zarzavat
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
zarzavat
4d ago
The hand wringing is premature. Thus far AI has only shown a superhuman aptitude for brute forcing proof of existence: - Disproof of the Jacobian conjecture by example - Construction of a non-sofic group - Existence of singularity in
2.
▲
by
zarzavat
6d ago
Because the complete power of Python also includes the power to fuck things up.
3.
▲
by
zarzavat
6d ago
Even if you judge OpenAI solely on their public communications it still sounds really bad. That they heard a rumour that a major open problem had been solved , so they decided to try and scoop the other mathematicians while they were wri
4.
▲
by
zarzavat
6d ago
It's more fundamental than that. AlphaZero is a shallower search with a heavier evaluation function. Stockfish is a deeper search with a lighter evaluation function. In chess, depth usually wins because of how narrow the search tree is
5.
▲
by
zarzavat
7d ago
The marginal improvements add up over time. Most people upgrade every few years, amortized over a 3 year upgrade cycle it's not really that expensive.
6.
▲
by
zarzavat
7d ago
Many mathematicians would be willing to end their careers for $1m. What makes this so sad is that OpenAI spent more than that for this empty PR stunt.
7.
▲
by
zarzavat
8d ago
Many people don't like the GPL for exactly that reason. Free software (i.e. copyleft) benefits from copyright at the expense of the wider open source community. Were it not for copyright then BSDs could take code from Linux and perhaps
8.
▲
by
zarzavat
8d ago
AI just solved a millennium prize problem. In a matter of days. Because of a rumor that someone else solved the same problem with AI. What exactly would AI have to do in order to not be called a bubble?
9.
▲
by
zarzavat
8d ago
It also fragments power, which makes it worth it.
10.
▲
by
zarzavat
10d ago
I agree and I was being facetious since I know the answer is nothing much. The point is that IBM write all this breathless marketing copy that seems detached from the actual utility of the device. Although I believe there's a differenc
11.
▲
by
zarzavat
11d ago
What's the largest semiprime it can factor?
12.
▲
by
zarzavat
11d ago
As the recent "proof" of the Collatz conjecture shows, that's not enough in an adversarial context. Human mathematicians don't submit proofs that take advantage of soundness bugs in Lean. AIs do.
13.
▲
by
zarzavat
11d ago
Are you talking about Japan? Because when I think of "Asia" strong job protections are just about at the bottom of the list.
14.
▲
by
zarzavat
12d ago
Humans are evolved to survive in the wild. We are not evolved for circuit design. Yet we can design circuits because evolution found it easier to develop a general problem solving nervous system than a nervous system which is adapted for ev
15.
▲
by
zarzavat
13d ago
> Humanity's Last Exam (w/ tools) This is one of the only benchmarks that actually matters for testing the frontier however. Other benchmarks can be gamed by simply being more persistent, but HLE is a diverse set of open-ended
16.
▲
by
zarzavat
13d ago
Therefore: the current transformer architecture is fundamentally incapable of AGI because the models have no mutable long-term memory. You only have weights (large immutable memory), or context (small mutable memory). Humans have mutable lo
17.
▲
by
zarzavat
14d ago
You jest but generating SVGs requires understanding of color, size, placement. It's stress testing visual/spatial/artistic capabilities that would be required for writing CSS/design work. Yes if you're doing backend
18.
▲
by
zarzavat
14d ago
You can see immediately that it's vibecoded so they kind of don't have a choice but to say so. It is interesting to see what can be achieved with AI even if few people want to play 100% vibecoded games.
19.
▲
by
zarzavat
15d ago
I'm polite to LLMs. It's not for the models it's for myself. If I start being rude to models then I might accidentally start being rude to other people as well.
20.
▲
by
zarzavat
19d ago
> This time it's a reCAPTCHA successor, which requires the user to scan a QR code using a Google/Apple owned phone. I can't wait for the future where I need a burner phone and a robotic arm just to browse the internet.
21.
▲
by
zarzavat
20d ago
Pizza dough is usually not rolled, it's stretched. If you roll it then you will press out the air bubbles which makes the final crust denser. That's not always a bad thing and some pizza makers do swear by rolling, but it does mea
22.
▲
by
zarzavat
20d ago
It's easy to automate making shitty pizzas (Pizza Hut, etc), but real pizzas require quite some dexterity and it seems it would be difficult to automate. Perhaps start with burgers, that would be easier.
23.
▲
by
zarzavat
20d ago
Chess is a brute force search problem. Humans are not good at chess, even a small computer can beat Magnus Carlsen. It would be better to compare models at how well they can write the code for chess engines, otherwise it's just sayin
24.
▲
by
zarzavat
21d ago
The objective of RL changed circa 2024. The 3.5-4o era models were trained by RLHF primarily to write in a way that's pleasing to humans. Starting with o1 the focus switched to reasoning, coding and benchmarks. If you remember when GPT
25.
▲
by
zarzavat
21d ago
If you spend all day talking to Claude then its mannerisms begin to grate on you like fingernails on a chalkboard. There's nothing wrong with the English per se, it's grammatically correct, it just sounds so "Claude". Cl
26.
▲
by
zarzavat
21d ago
I thought 'how bad can it be?', yet only three paragraphs in "a specifier now names a repository rather than choosing a transport" If you're going to serve us slop then at least serve us good slop, give it a pas
27.
▲
by
zarzavat
21d ago
Even Artificial Analysis has Opus 5 better than Fable in their aggregated "Intelligence Index" which combines 9 benchmarks. Opus 5 is heavily benchmaxxed.
28.
▲
by
zarzavat
21d ago
It's China. It's a given that they use your data for training. At least they're nice enough to be honest about it.
29.
▲
by
zarzavat
22d ago
3. Deployment problems unrelated to the weights causing degraded performance
30.
▲
by
zarzavat
22d ago
Freudian slip.
More ›