Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
data_maan
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
data_maan
1mo ago
It always amazes me how a random dump of someone who read the first 40 pages of PM attracts dozens comments on HN. This really must be a very math-starved community of people who wanted to learn math but never quite could.
2.
▲
by
data_maan
3mo ago
Makes me wonder what the value of a human on a PhD course is
3.
▲
by
data_maan
3mo ago
Michal Valko's paper should be mentioned. It's discussed here https://youtu.be/UUq4ixTmye8?is=K75EsFIYgyrPKdCI
4.
▲
by
data_maan
3mo ago
Another famous dude dumping his thoughts on HN who is gulping it up like an addict. Add this to the long list of names like Terence Tao, and others who seem to be intellectually incontinent lately in the sense that one cannot navigate this
5.
▲
by
data_maan
3mo ago
Aside from the sad life events, little information is shares about his "system", the thing HN is interested in. Is this more than a harness built on top of a SOTA commercial LLM?
6.
▲
by
data_maan
3mo ago
Was this in the GPT2 paper?
7.
▲
by
data_maan
5mo ago
If LLMs lie as much as the OP claims in the article, why can they then solve Olympiad math problems they never saw during training, consistently? There's the aimoprize.com on Kaggle for example that shows this
8.
▲
by
data_maan
6mo ago
More "American minds": https://en.wikipedia.org/wiki/Hartmut_Esslinger Chief designer at Apple war German.
9.
▲
by
data_maan
6mo ago
> built with American capital and mostly American minds. I would say "built with American agency and commercial spirit", not minds. Most of the things that we have were first built elsewhere (Germany being a prime supplier here
10.
▲
by
data_maan
6mo ago
To be fair, Iran is not pretentious either, killing a few thousand people because they dared to protest. There are no good guys in this conflict.
11.
▲
by
data_maan
6mo ago
https://www.worldatlas.com/us-history/wars-the-united-states...
12.
▲
by
data_maan
6mo ago
A model to whose internals we don't have access solved a problem we didn't knew was in their datasets. Great, I'm impressed
13.
▲
by
data_maan
6mo ago
Strategic? Yes. Moral? Hm. From a moral POV this would be about who has the right to terrorize the Iranian population: the Iranian government or the US/Israel government.
14.
▲
by
data_maan
6mo ago
Opinions differ: hobby coders love it, but domain expert secretly despise it because it narrows the gap between the skills they spent years honing and the average Claude, I mean Joe, that just uses this mental exoskeleton.
15.
▲
by
data_maan
6mo ago
> The people getting pushed out are the intermediates and seniors who aren't high performers. Also the people that can't market themselves. There are very average programmers that have a large following on X that seem to do ver
16.
▲
by
data_maan
7mo ago
I love these posts that are so on the edge that I can't tell if it's sarcastic or for real :)
17.
▲
by
data_maan
7mo ago
> What do you mean ? These are top-notch mathematicians YeS. I didn't dispute that. I disputed that they are NOT top notch ML specialist and have made one of the worst benchmarks of 2025-2026. Benchmarks like these would have worked
18.
▲
by
data_maan
7mo ago
If it's the latter case (which it has to be), it seems that attention credit (via, e.g., articles in NY Times) is very unfairly distributed. None of the people that advanced the state of benchmarking and did the hard work on much bigge
19.
▲
by
data_maan
7mo ago
> We will learn if the magical capabilities attributed to these tools are really true or not. They're not. We already know that. FrontierMath. Yu Tsumura's 553th problem, RealMath benchmark. The list goes on. As I said many tim
20.
▲
by
data_maan
7mo ago
> These problems are representative of the types of subproblems research mathematicians have to solve to get a “research result”. They are finding that LLMs aren’t that useful for mathematical research because they can’t crush these prob
21.
▲
by
data_maan
7mo ago
But everything has been explored in other datasets already. If only a bunch of mathematicians learn something, why are so many people talking about this, why is the NY Times posting about this? This is the attention economy at its worst.
22.
▲
by
data_maan
7mo ago
It's not angst. It's intense frustration that they 1) are not doing the science correctly, and 2) that others (e.g. FrontierMath) already did everything they claim to be doing, so we won't learn anything new here, but somehow
23.
▲
by
data_maan
7mo ago
If you want to do this rigorously, you should run it as a competition like the guys at the AI-MO Prize are doing on Kaggle. That way you get all the necessary data. I still think this is bro science.
24.
▲
by
data_maan
7mo ago
> There are some experiments which cannot be carried out more than once Yes, in which case a very detailed methodology is required: which hardware, runtimes, token counts etc. This does none of that.
25.
▲
by
data_maan
7mo ago
It wasn't like this in any way. CASP relies on a robust benchmark (not just 10 random proteins), and has clear participation criteria, objective metrics how the eval plays out, etc. So I stand by my claim: This isn't scientific. I
26.
▲
by
data_maan
7mo ago
Yes, but people at those labs may be running those problems because a Fields Medalist is in the paper, and it got hype. Not because of the problems, and not because this is new methodology. And once the labs report back, what do we know tha
27.
▲
by
data_maan
7mo ago
Because the companies have the data and can solve them -- so providing the question to a company with the necessary manpower, one cannot guarantee anymore that the solution is not known, and not contained in the training sample.
28.
▲
by
data_maan
7mo ago
How is that interesting for a scientific point of view? This seems more like a social experiment dressed as science. Science should be about reproducibility, and almost nothing here is reproducible.
29.
▲
by
data_maan
7mo ago
On the website https://1stproof.org/#about they claim: "This project represents our preliminary efforts to develop an objective and realistic methodology for assessing the capabilities of AI systems to autonomously sol
30.
▲
by
data_maan
7mo ago
> these are problems of some practical interest, not just performative/competitive maths. FrontierMath did this a year ago. Where is the novelty here? > a solution is known, but is guaranteed to not be in the training set for any
More ›