Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
adroniser
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
adroniser
5mo ago
what are you even yapping about.
2.
▲
by
adroniser
8mo ago
Adding the position vector is basic sure, but it's naive to think the model doesn't develop its own positional system bootstrapping on top of the barebones one.
3.
▲
by
adroniser
1y ago
fmri's are correlational nonsense (see Brainwashed, for example) and so are any "model introspection" tools.
4.
▲
by
adroniser
1y ago
Of the papers submitted to a conference, it might be that reviewers don't offer suggestions that would significantly improve the quality of the work. Indeed the quality of reviews has gone down significantly in recent years. But if Ant
5.
▲
by
adroniser
1y ago
So you think that this blog post would make it into any of the mainstream conferences? I doubt it.
6.
▲
by
adroniser
1y ago
peer review would encourage less hand wavy language and more precise claims. They would penalize the authors for bringing up bizarre analogies to physics concepts for seemingly no reason. They would criticize the fact that they spend the wh
7.
▲
by
adroniser
1y ago
didn't hackers used to be for piracy?
8.
▲
by
adroniser
1y ago
This suggests people should pre-register benchmarks. Because currently it feels like there is little incentive to publish benchmarks that models saturate.
9.
▲
by
adroniser
2y ago
But there are lots of models available now that render much faster which are better quality than sora
10.
▲
by
adroniser
2y ago
I completely agree this shit is so depressing. When I saw the AlphaProof paper I basically spent 3 days in mourning basically, because their approach was so simple.
11.
▲
by
adroniser
2y ago
I think the whole paper is a satire lol.
12.
▲
by
adroniser
2y ago
Does it really? If you want an LLM to edit code you need to feed it every single line of code in a prompt. Is it really that surprising that having just learnt it has been timed out, and then seeing code that has an explicit timeout in it,
13.
▲
by
adroniser
2y ago
I agree with you that transformers are probably not the architecture of choice. Not sure what that has to do with the viability of RL though.
14.
▲
by
adroniser
2y ago
Hmm well the reason a pre-trained transformer is a fancy sentence completion engine is because that is what it is trained on, cross entropy loss on next token prediction. As I say, if you train an LLM to do math proofs, it learns to solve 4
15.
▲
by
adroniser
2y ago
But RL algorithms do implement things like curiosity to drive exploration?? https://arxiv.org/pdf/1810.12894 . Thinking to arbitrary depth sounds like Monte Carlo tree search? Which is often implemented in conjunction w
16.
▲
by
adroniser
2y ago
Isn't RL the algorithm we want basically?
17.
▲
by
adroniser
2y ago
The distinction is that LLMs are not used for what they are trained for in this case. In the vast majority of cases someone using an LLM is not interested in what some mixture of openai employees ratings + average person would say about a t
18.
▲
by
adroniser
2y ago
How about you want to solve sudoku say.And you simply specify that you want the output to have unique numbers in each row, unique numbers in each column, and no unique number in any 3x3 grid. I feel like this is a very different type of pro
19.
▲
by
adroniser
2y ago
"AlphaProof is a system that trains itself to prove mathematical statements in the formal language Lean. It couples a pre-trained language model with the AlphaZero reinforcement learning algorithm, which previously taught itself how to
20.
▲
by
adroniser
2y ago
I did read your entire comment, and that is what prompted my response, because from my perspective your entire premise was based on LLMs failing at simple examples, and yet despite admitting you thought there was a chance an LLM would succe
21.
▲
by
adroniser
2y ago
If you're going to suggest something you think an LLM can't do I think at the very least as a show of good faith you should try it out. I've lost count of the number of times people have told me LLMs can't do shit that t
22.
▲
by
adroniser
2y ago
Yes but you are not taking an uncountable union. You are taking a finite union.
23.
▲
by
adroniser
2y ago
To be clear, the construction given here violates the finite additivity property of measure. It's got nothing to do with the countable/uncountable additivity property.
24.
▲
by
adroniser
2y ago
Essentially the difficulty arises from attempting to assign a measure (area) to every single subset of the sphere, where you say that rotations need to preserve this measure. The paradox can be viewed as a proof that you cannot assign a mea
25.
▲
by
adroniser
2y ago
The only source I can find for this estimate is from a year ago. I feel like efficiency has gone up by a lot since then
26.
▲
by
adroniser
2y ago
anabolic steroids will kill you idk why you'd want to mess with them.
27.
▲
by
adroniser
2y ago
And yet it doesn't rule out that it can't. See new york times lawsuit
28.
▲
by
adroniser
2y ago
I don't see how the point about the typical human is relevant. Either you can reason or you can't, the ARC test is supposed to be an objective way to measure this. Clearly a vanilla LLM currently cannot do this, and somehow an exp
29.
▲
by
adroniser
2y ago
The curious thing is if you can ever hit a snowball type point, because by proving theorems you are generating training data, and maybe you can get to a point where you effectively have limitless data.
30.
▲
by
adroniser
2y ago
Doubtful. Apple can get away with tiny iterations on smartphones because they have the brand and they know people will always buy their latest product. LLMs aren't physical products so there is no cost to switching other than increased
More ›