Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rsfern
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
rsfern
7d ago
I agree (and so does Buckmaster based on his written statement) that we are better having solved this. But I disagree that which humans were credited is the heart of the issue in this particular controversy. The question is what do you need
2.
▲
by
rsfern
7d ago
That’s the prevailing narrative, but I think this controversy calls it into question to some extent. If the OpenAI result wouldn’t have been possible without experts seeding the training data with feedback on promising solution routes, ther
3.
▲
by
rsfern
7d ago
I think you’re missing an important distinction. “Major damage” to the talent pipeline because models become capable of original end-to-end mathematics is what the community has been discussing. But if the models rely on sniping nearly comp
4.
▲
by
rsfern
7d ago
The session data could be cryptographically signed. Probably easier in an open harness?
5.
▲
by
rsfern
7d ago
The screen cap of their full correspondence in the Twitter thread, which was the subject of the preceding sentence. The Twitter post has an obviously incomplete fragment of the conversation that doesn’t resolve what the author presents it a
6.
▲
by
rsfern
7d ago
Without seeing the full correspondence it’s hard to evaluate for sure, but parent linked to a tweet from the OpenAI employee at the center of the controversy, that’s a primary source you can read and evaluate yourself Personally I don’t fin
7.
▲
by
rsfern
7d ago
It does, yes. So designing objections functions and making sure you can afford the training rollouts becomes really important in defining which problems are tractable. It will be really interesting to see how that shapes the kinds of proble
8.
▲
by
rsfern
8d ago
Agreed, but i think this underscores my point. We have numerical simulations in materials science too, but that doesn’t mean formally verified theorems about the underlying equations automatically translate to formal (or even informal) veri
9.
▲
by
rsfern
8d ago
Yes, definitely! There’s a long history of this and I think there’s tons of opportunities for more. Both for improving the exactness/physical fidelity of models and for developing new approximate theories and simulation methods
10.
▲
by
rsfern
8d ago
Regardless of what you think of the priority dispute issue discussed on sibling threads, I’m highly skeptical of the closing quote that this Navier Stokes result means that the same approach of casually spending a few million on agentic com
11.
▲
by
rsfern
8d ago
Thanks! This seems really cool. If I’ve got it right, your UI builds and displays these diffs (and implements undo/redo) by parsing the edit tool calls? If so that seems really nice, one of the things I don’t like with coding agents is
12.
▲
by
rsfern
8d ago
Interesting project idea! The link seems to be 404, is the repo still private?
13.
▲
by
rsfern
8d ago
Why would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect mo
14.
▲
by
rsfern
1mo ago
The bit you quoted doesn’t capture why the reporters are upset: > White House journalists are outraged that a threat credible enough to force Donald Trump to escape from Air Force One using an airport catering truck wasn’t relayed to th
15.
▲
by
rsfern
1mo ago
Right, I did specifically say that most of the DOE scientists are contractors, but I concede the phrase “government scientist” is a bit ambiguous. I appreciate the extra detail you added. I think the distinction between political appointee
16.
▲
by
rsfern
1mo ago
Let’s distinguish a bit. There are political appointees (Trump’s government employees as you say) who are mostly upper management, and there are career civil servants (all the government scientists are under this category) who have a strong
17.
▲
by
rsfern
1mo ago
I found this paper really thought provoking, but I think the conclusion of “world models are the solution” leaves something to be desired. People are already equipping agentic systems with physical simulation tools and exploring action-cond
18.
▲
by
rsfern
2mo ago
Or they could store the reading traces and validate the user hasn’t edited them server-side? They could sign reasoning traces so they can’t be counterfeited?
19.
▲
by
rsfern
2mo ago
The cell DAG enforces that there’s no implicit state, which reduces cognitive load for me a lot and provides some pressure to abstract experimental code into functions. In Jupyter this is left to user discipline and restart-and-run-all work
20.
▲
by
rsfern
2mo ago
My point with the force field example wasn’t to argue against neural scaling as a valid strategy, it totally is effective and a lot of groups are doing it. But I feel like we might be talking past each other a bit. What I’m pushing back on
21.
▲
by
rsfern
2mo ago
I don’t think there’s a fundamental reason that performance has to be monotonic in model size or even training FLOPs. At least I don’t think it’s been proved to be so, so I think “misinformed” is a bit premature and sort of makes GP’s point
22.
▲
by
rsfern
2mo ago
I don’t mean to pick on you in particular here, but this approach has been bothering me a lot lately, and it seems like it’s super common. I get that it’s an early prototype and not all the design choices are made yet, but I struggle with “
23.
▲
by
rsfern
2mo ago
I often start prompts with “please”, but I usually don’t thank the model. Framing a question or a request for help with “please” is in distribution for me, it’s a distraction from composing a thoughtful prompt about my actual question to go
24.
▲
by
rsfern
2mo ago
The DHS secretary seems to me to have the point of hosting international students backwards > This final rule ensures that foreign students remain focused on their primary purpose: completing their studies and returning home.” Especial
25.
▲
by
rsfern
2mo ago
True. But it has no idea that it has no idea, so it might be able to look back at the session trace and pattern match its way to actionable feedback?
26.
▲
by
rsfern
2mo ago
I think JEPA is super interesting, but I feel like this example highlights some of the challenges of long horizon planning. For one, chunking the planning stage into a bunch of intermediate goals seems really limiting, because a lot of what
27.
▲
by
rsfern
2mo ago
What aspect do you consider basic? I haven’t had a chance to read more than the abstract because of the paywall, but the really interesting thing here is the mechanically induced transformation that leads to a 3-phase nanocrystalline alloy.
28.
▲
by
rsfern
2mo ago
The science daily article is just incorrect to call this a superalloy, which it is not. This is a high entropy refractory alloy (HfNbTaTiZr), superalloys are usually based on lighter metals and they usually have only one dominant element w
29.
▲
by
rsfern
2mo ago
This is really cool metallurgy. They start with an alloy and deform it and because of elemental size mismatch they can cause the alloy to self assemble into nanoscale crystals with three different structures The paper: https://ww
30.
▲
by
rsfern
2mo ago
Terry Tao has actually been one of the more prominent voices in the math community exploring AI for cutting edge mathematical discovery. This particular post is a bit softer but he has also written a lot about using AI assistance for seriou
More ›