Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
t55
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
Sokoban Speedrun for RL
(github.com)
6 points
by
t55
2mo ago
|
0 comments
2.
▲
by
t55
2mo ago
what timelines do you have in mind
3.
▲
RL Speedrun
(github.com)
2 points
by
t55
3mo ago
|
0 comments
4.
▲
Target Policy Optimization
(arxiv.org)
1 points
by
t55
5mo ago
|
0 comments
5.
▲
Show HN: Kilroy – Knowledge base for teams using Claude Code
(github.com)
5 points
by
t55
5mo ago
|
0 comments
6.
▲
Procedural Reasoning Datasets
(github.com)
1 points
by
t55
1y ago
|
0 comments
7.
▲
In Defence of Gary Marcus
(reubenadams.substack.com)
3 points
by
t55
1y ago
|
0 comments
8.
▲
Reasoning Gym – Procedural RL reasoning datasets
(github.com)
1 points
by
t55
1y ago
|
0 comments
9.
▲
ChatGPT Agent [video]
(youtube.com)
3 points
by
t55
1y ago
|
0 comments
10.
▲
by
t55
1y ago
that's a standard feature in cursor, windsurf, etc.
11.
▲
by
t55
1y ago
this is what deepmind did 10 years ago lol
12.
▲
by
t55
1y ago
For a 100k token context window; all those models are comparable though gemini 2.5 pro shines for 200k+ tokens
13.
▲
by
t55
1y ago
i didn't say they invented everything; in science you always stand on the shoulders of giants i still think my original statement is fair
14.
▲
by
t55
1y ago
yeah, RLVR is still nascent and hence there's lots of noise. > How can these spurious rewards possibly work? Can we get similar gains on other models with broken rewards? it's because in those cases, RLVR merely elicits the rea
15.
▲
by
t55
1y ago
so you think it's fake news? another example of a paper with strong claims without much evidence?
16.
▲
by
t55
1y ago
agree, the RG evals feel like a fresh breeze
17.
▲
by
t55
1y ago
> prolonged RL training can uncover novel reasoning strategies that are inaccessible to base models, even under extensive sampling does this mean that previous RL papers claiming the opposite were possibly bottlenecked by small datasets?
18.
▲
by
t55
1y ago
> I personally think that Gemini 2.5 Pro's superiority comes from having hundreds or thousands RL tasks (without any proof whatsoever, so rather a feeling). Given that GDM pioneered RL, that's a reasonable assumption
19.
▲
by
t55
1y ago
it aged well!
20.
▲
ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
(arxiv.org)
105 points
by
t55
1y ago
|
28 comments
21.
▲
Show HN: Rehearsal.so, Duolingo for Public Speaking
(rehearsal.so)
3 points
by
t55
1y ago
|
1 comments
22.
▲
by
t55
1y ago
Very cool, reminds me of https://rehearsal.so/ which sort of interviews will you support?
23.
▲
End-to-End Vision Tokenizer Tuning
(arxiv.org)
3 points
by
t55
1y ago
|
0 comments
24.
▲
YC Interview Mock Practice
(rehearsal.so)
2 points
by
t55
1y ago
|
0 comments
25.
▲
D1: Scaling Reasoning in Diffusion LLMs via Reinforcement Learning
(dllm-reasoning.github.io)
4 points
by
t55
1y ago
|
0 comments
26.
▲
Are LLMs more than autocomplete? AI Debate
(rehearsal.so)
1 points
by
t55
1y ago
|
0 comments
27.
▲
Block Diffusion: Interpolating Autoregressive and Diffusion Language Models
(m-arriola.com)
72 points
by
t55
1y ago
|
16 comments
28.
▲
How to stay in flow while using Cursor or Windsurf
(rehearsal.so)
2 points
by
t55
1y ago
|
0 comments
29.
▲
by
t55
1y ago
great article!
30.
▲
Generative Modelling in Latent Space
(sander.ai)
2 points
by
t55
1y ago
|
0 comments
More ›