Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
danielmarkbruce
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
danielmarkbruce
22h ago
Nobody is suggesting nobody or group cooperates on anything ever. This comment could go in the definition section of "strawman".
2.
▲
by
danielmarkbruce
5d ago
and taken to it's extreme.... isn't the atmosphere secretly taking the vibrations in the air right outside out mouth and propagating them through a physical medium, making them available to anyone to capture?
3.
▲
by
danielmarkbruce
5d ago
Yeah, good catch. I was under the impression they did the proof in lean from the get go, but you are right. I guess the nature of the problem lent itself to the 10k agents. Ie, there isn't something general to take here.
4.
▲
by
danielmarkbruce
5d ago
You are hitting the google model through openrouter or directly?
5.
▲
by
danielmarkbruce
6d ago
There might not be a good abstraction. I've built a few harnesses for different types of workflows, and the details are so different I struggle to see a good abstraction. It's also not clear there should be - if you look at most c
6.
▲
by
danielmarkbruce
7d ago
Yup, and maybe the law will need to evolve. Fwiw, imo, you should assume everything you say (and write) is recorded going forward.
7.
▲
by
danielmarkbruce
7d ago
People generally don't think of keeping bits in memory for some short period of time as "recording".
8.
▲
by
danielmarkbruce
7d ago
From a legal perspective you can't record something in many us states. Using the microphone, keeping the bits in memory while you turn it into text, probably doesn't meet any reasonable definition of "record".
9.
▲
by
danielmarkbruce
7d ago
You've always been allowed to transcribe (ie, write) out a conversation. Recording is a different thing. A person can always claim, easily and believably, that the transcription is made up. It's just text.
10.
▲
by
danielmarkbruce
7d ago
Agree, it is effectively understanding human biology - which we suck at, have little data on, have few good models for. So, it doesn't appear there is any way to avoid the clinical trial process. And it won't speed up, and it won&
11.
▲
by
danielmarkbruce
8d ago
It's closer to strategizing once you've run a model through RL. It's optimized to take steps which will lead it to a good outcome down the road.
12.
▲
by
danielmarkbruce
8d ago
Well, amazon and apple haven't done great either. One might reasonably claim they weren't as well placed as goog or msft, but it might just be "big company can't do genuinely new thing".
13.
▲
by
danielmarkbruce
8d ago
No, it won't. How do you verify some causal claim in biology? The reason AI is doing so well in math proof writing is that it can verify every idea it has, quickly.
14.
▲
by
danielmarkbruce
8d ago
Maybe. Maybe not. Look at AI drug design - it's not really speeding up the important part - drug trials. There isn't really a coherent plan to use AI for the most complex part of drug discovery at all.
15.
▲
by
danielmarkbruce
8d ago
If you can create a graph of independent work, which you can with many such problems, agents can work together nicely. Again, thank Lean and the tooling around it.
16.
▲
by
danielmarkbruce
8d ago
Highly capable of writing math proofs, no doubt. It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue
17.
▲
by
danielmarkbruce
8d ago
If you look closely at the gains in math, it's largely in proof writing. The reason is Lean, it's not some general intelligence jump, and the number of people actually working on proofs in life rounds to zero.
18.
▲
by
danielmarkbruce
8d ago
With Lean, math has become a really well suited problem for LLMs. We will likely see large gains for many years from here, just doing more and more rlvr, like continuously, non stop. No need to train from scratch. It really doesn't sp
19.
▲
by
danielmarkbruce
9d ago
Ok. This all makes sense I guess. Good going, hope it goes well.
20.
▲
by
danielmarkbruce
9d ago
It's more responsible than the use of electricity for a messaging board for people to argue minutiae.
21.
▲
by
danielmarkbruce
9d ago
Why not specify 2-3 problems? Won't you have people show up having already spent a bunch of time on their self chosen problem? Which sort of defeats the point of seeing what you can do in a short period of time?
22.
▲
by
danielmarkbruce
10d ago
But the next move is self evident if you have a prediction of the value of being in each of the states possible.
23.
▲
by
danielmarkbruce
10d ago
Your initial comment says "reward". Reward and return are not the same thing. The policy is choosing the moves based off of returns at the next state, not the immediate rewards. One choice might have reward 0 and expected return 1
24.
▲
by
danielmarkbruce
11d ago
I'm not the one hiding behind a fake name. If you want to understand how this stuff works, there are totally decent books about building them from scratch. It's not that hard, and you'll likely find it interesting. Sebastian
25.
▲
by
danielmarkbruce
11d ago
When you are done with the section on RLVR, consider whether the model is predicting tokens, or making moves. There is a reason the word "policy" is used in RL.
26.
▲
by
danielmarkbruce
11d ago
Many people are in fact claiming the thing you are saying they are not - even if you are not. The reason they are claiming it is that it was true at one point, and most intro courses/blog posts/videos still describe them that way
27.
▲
by
danielmarkbruce
11d ago
Brush up :) The policy is optimized to maximize the the total reward, defined as the sum of the reward at each step, discounted by some factor.
28.
▲
by
danielmarkbruce
11d ago
Mine was sarcasm. People who actually understand cars have built them. Until you build something, you don't understand it.
29.
▲
by
danielmarkbruce
11d ago
Nathan Lambert wrote a good book recently, and he and his team wrote the paper below about Tulu 3 (Allen Institute). Both are good reads. https://arxiv.org/pdf/2411.15124
30.
▲
by
danielmarkbruce
11d ago
The policy is not learned token by token during RLHF and RLVR. The reward model doesn't score token by token.
More ›