Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
hodgehog11
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
hodgehog11
8d ago
This is an insane thing to read. Bubeck had a reputation even before he started with OpenAI. Of course it was him that was involved in this drama. This is such a sad mess, and it really didn't have to be this way.
2.
▲
by
hodgehog11
8d ago
It's literally a toggle in the options for ChatGPT, one which is on by default and most researchers probably have on without realising it. So to say that it is unlikely is extremely suspicious. No, they did not literally pull user da
3.
▲
by
hodgehog11
8d ago
Yes, mathematicians are. And yes, most of my colleagues did not even know the opt-out was an option.
4.
▲
by
hodgehog11
10d ago
Touche, I used to work in pure probability where that ridiculous Hardy-Littlewood rule used to cause all sorts of problems, but now work in statistics, where it is no longer an issue. To be clear, the colleagues I am referring to mostly wor
5.
▲
by
hodgehog11
10d ago
Uh, care to explain? I have several colleagues that stopped submitting to journals once they reached full professor. They only submit papers from their students for the benefit of their careers. First-author papers, not so much.
6.
▲
by
hodgehog11
11d ago
Exactly, and the advantage is that checking that the problem is "formalized" here is essentially isolated to verifying that the final theorem statement matches the claim. If there are no 'sorry's and the program compiles
7.
▲
by
hodgehog11
11d ago
You can prove that doing this will spiral training into a fixed point. There was a lot of research into getting this to work in the past, but it never truly worked well. The hope was that if RLVR was used quite a bit, and the general perfor
8.
▲
by
hodgehog11
11d ago
Why would they send it out for "expert review"? Every time, they have just made the AI generate a Lean proof. In fact, it seems like the most plausible direction to NS is computationally assisted detection of a blowup solution, wh
9.
▲
by
hodgehog11
11d ago
It basically is a formality at this level. Many top math researchers now hardly even submit to journals at all and just put up a preprint. At this scale, peer review happens by the audience. They don't need a journal to get people re
10.
▲
by
hodgehog11
16d ago
I also believe this. Post-training LLMs with vague metrics can only be achieved with RLHF, which is not impossible, but extremely costly and difficult. Instead, companies will opt for RLVR, focusing on math and programming tasks. This pushe
11.
▲
by
hodgehog11
17d ago
I agree that this should be something that researchers reflect on. GPT-2 is one of the primary models to research on nowadays, and many recent developments have come from studying it as a test bench. Imagine if CRISPR was considered "t
12.
▲
by
hodgehog11
22d ago
It's good to see validated numerical proofs seeing a resurgence now that they are substantially easier to achieve. Others might be able to chime in, but my experience is that AI is effectively taking proofs that were once iterative (pu
13.
▲
by
hodgehog11
24d ago
In a frontier scientific research environment, funding is often limited, so personal subscriptions are more common. Fable can hit a 5-hour usage limit on the Max subscription tier before it finishes a single complex math prompt . Most of t
14.
▲
by
hodgehog11
26d ago
I hope you understand the context in which that was said. The point of that statement is that the only way to rigorously verify correctness of a program is by using formal methods. Those are often too difficult to achieve by humans, which i
15.
▲
by
hodgehog11
26d ago
No it really is about the test suite, and provably so. As another poster pointed out, speed is a superoptimization problem and the test suite provides the constraints. If the constraints are appropriately set, even a naive genetic algorithm
16.
▲
by
hodgehog11
1mo ago
The Chinese labs are picking up on the low hanging fruits on efficiency, and no, you do not need to abandon transformers, you just need to push them closer to the more computationally efficient architectures of the past. OpenAI and Google
17.
▲
by
hodgehog11
1mo ago
> The game often has its own DRM though which will stop you I think you missed the "know where to look" part. It's called a Steam emulator, for starters. Note that I speak about this strictly for preservation purposes, as
18.
▲
by
hodgehog11
1mo ago
Even if you get all of your games via Steam, provided you have them downloaded, you can still run them without Steam if you know where to look. Obviously GOG is far better in this regard, but preservation is not a concern on PC, outside of
19.
▲
by
hodgehog11
2mo ago
My expertise lies in deep learning theory, and yes, the "intelligence" is coming primarily from scaling up, among other things. There are good reasons for this, but essentially it comes down to taking advantage of a narrow statist
20.
▲
by
hodgehog11
2mo ago
Agreed, AI is not capable at the moment of coming up with radical ideas to solve the tough problems. Sadly, I would argue many problems in math are likely to be found to be not actually tough in this sense, and those working in "comfor
21.
▲
by
hodgehog11
2mo ago
Neither Claude nor GPT are acceptable for writing English text. Personally I have found Gemini to be far better, and that is really all I use it for.
22.
▲
by
hodgehog11
2mo ago
That has always been the major strength of GPT, that's the model you use for checking. It often nearly isn't as good for creation though.
23.
▲
by
hodgehog11
2mo ago
Agreed. The benchmark closest to my experience is FrontierMath Tier 4. Fable and Sol (90%) are very far ahead of Kimi K3 (not even 40%). Kimi is trained heavily to basic agentic tasks, like all the other open models right now.
24.
▲
by
hodgehog11
2mo ago
They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, which is what most people use these models for.
25.
▲
by
hodgehog11
2mo ago
It really does depend on your application. In my domain (math research), it is substantially better. Fable can solve really hard tasks with surprising consistency. It makes mistakes, and occasionally refuses, but honestly, at the top level,
26.
▲
by
hodgehog11
2mo ago
Tell that to my colleagues. Despite Sol getting the attention, Fable is really starting to have an impact on mathematicians right now. It has unbelievable insights in a lot of cases that can rapidly speed up progress.
27.
▲
by
hodgehog11
2mo ago
I think it is pretty safe to say at this point that having large open LLM models available is better for humanity than them remaining proprietary. Echoing Linus Torvalds' recent comments, AI is genuinely useful right now, and is here t
28.
▲
by
hodgehog11
2mo ago
The "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that m
29.
▲
by
hodgehog11
2mo ago
That would go against everything that Dario believes in (note that I refer to the CEO and not the company; the staff at Anthropic are not so ridiculous). He believes in Anthropic being the sole arbiter of the forefront of this technology, b
30.
▲
by
hodgehog11
2mo ago
Value models are always going to be there; you can always distill from a larger model. Having a really intelligent model, regardless of the size, is much better for building confidence in your brand. That is a big reason why the US companie
More ›