Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kzrdude
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
kzrdude
3d ago
This is tricky, because we really want language-independent training of skills. We know that self-play type of reinforcement learning is incredibly effective when possible. But at the same time, they are our tools - so we need supervised la
2.
▲
by
kzrdude
3d ago
We need to recognize this as a failure in training. It did some useful stuff but it can be much better. A training signal is likely missing.
3.
▲
by
kzrdude
4d ago
Not only in America
4.
▲
by
kzrdude
4d ago
But if I remember correctly, he gained recognition for his achievement rather quickly after posting.
5.
▲
by
kzrdude
4d ago
If we read the link, it has a section called Gold Standard: comparator and external checkers , and comparator is how OpenAI has gone about checking their lean proofs.
6.
▲
by
kzrdude
5d ago
I wonder what Demis Hassabis thinks about this. I thought he cared a lot about mathematics.
7.
▲
by
kzrdude
5d ago
You should read his comment from wednesday here: https://terrytao.wordpress.com/2026/09/07/finite-time-blowup... And yes, at one point it gives you right about them taking advantage of him. He sat down for an
8.
▲
by
kzrdude
5d ago
Well OpenAI took one step on the back foot at least, withdrawing from sponsoring this math hackathon event https://xcancel.com/danintheory/status/2098125701782372640
9.
▲
by
kzrdude
5d ago
My university has an agreement with Microsoft copilot. We can log into copilot in many ways, and it's only if you log in the correct way that you get the "Enterprise Data Protection" copilot version, with a green shield symbo
10.
▲
by
kzrdude
5d ago
It seems like we're not supposed to care about the code quality then? I guess that's the compilers argument. But I'm not ready to give up the code just yet.. These LLMs don't even have a stable interface, they change eve
11.
▲
by
kzrdude
5d ago
They have plausible deniability on that one: not making any profit.
12.
▲
by
kzrdude
6d ago
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
13.
▲
by
kzrdude
6d ago
Interesting, and if you don't mind, where do we put humans (and superhumans like Magnus Carlsen)? I think they have a heavy evaluation function and do shallower search.
14.
▲
by
kzrdude
6d ago
There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback. That
15.
▲
by
kzrdude
6d ago
But the very fact that you go to "chatgpt.com" and write to them; "Dear Diary, today I thought.."; there is no reason they would not receive and process your data, unless explicitly promising not to (which also requires
16.
▲
by
kzrdude
6d ago
And just a day later we have the tech report available for V4.1: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
17.
▲
by
kzrdude
6d ago
Isn't one problem that it's hard to determine what equivalent compute is, for CPU search vs a neural net based engine like AZ or Leela?
18.
▲
by
kzrdude
6d ago
V4 Flash was one of the big events of this year, and its already retired and replaced by V4.1 Flash.
19.
▲
by
kzrdude
6d ago
The main claim, that "they do not seem to have any mathematicians capable of understanding what they put out", was also corroborated by Sebastien (OpenAI) who explained they don't have any experts on Navier-Stokes.
20.
▲
by
kzrdude
7d ago
The Chinese labs have released interesting papers to accompany their releases too, especially DeepSeek and Kimi. This improves their standing among a few of us, who really like to see and read the papers with details about what they have ch
21.
▲
by
kzrdude
7d ago
Now you are using the legislators' rationale that money is being made so we need to protect that, forgetting to weigh the cost for the community.
22.
▲
by
kzrdude
7d ago
It's got 'max' and 'ultra' but I don't see a setting for level-headed.
23.
▲
by
kzrdude
7d ago
On the contrary, a level-headed summary that gathers information from all the different sources is necessary.
24.
▲
by
kzrdude
7d ago
Depends on how much risk they are willing to accept. What is strange here is that it's clear that OpenAI is both a service provider and a competitor to mathematicians. It almost reminds me of Amazon which both hosts external merchants
25.
▲
by
kzrdude
8d ago
We come back to the rule: "The cloud is just someone else's computer". The way for people or companies or universities to control their data and information is to keep it on their own computers.
26.
▲
by
kzrdude
8d ago
Let's see what happens. In contrast to this whirlwind of math that's going on right now, the millennium prize rules require publishing in a reputable journal and 2 years of waiting time to establish that the proof has been accepte
27.
▲
by
kzrdude
8d ago
The construction is that there is one file you need read and verify, the challenge file. If you've verified that file and trust that your lean compiler works correctly, the proof will be correct. That file should be https://
28.
▲
by
kzrdude
8d ago
I don't see how it would be negative PR. If anything, the love these breakthroughs and use it in their PR campaigns.
29.
▲
by
kzrdude
8d ago
This "fefferman options c and d" thing sounds damning but that's nothing. Let's assume the forelaid proof is correct. Then option C or D is the only way to win the prize, those options are the only ones that solve it. Th
30.
▲
by
kzrdude
8d ago
If you don't trust the other party, then it doesn't matter how the checkbox is set. The fundamental rule, IMO, is don't send precious or secret data to a third party.
More ›