Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
imenani
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
14 ms
·
1.
▲
by
imenani
4d ago
The primary source https://darioamodei.com/post/we-must-pace-the-frontier#top
2.
▲
We Must Pace the Frontier
(darioamodei.com)
3 points
by
imenani
4d ago
|
0 comments
3.
▲
by
imenani
12d ago
I don’t think next token prediction is a particularly good description of pretraining either. The intermediate representations at each position are being optimised not only to help predict the next token, but also to help predict all subseq
4.
▲
by
imenani
1mo ago
Regarding subsidy/profits: Anthropic at least seems to be on track to report a profitable Q2 2026 > Anthropic has projected revenue of at least $10.9 billion for the second quarter of 2026, more than double the previous quarter, o
5.
▲
by
imenani
2mo ago
For anyone wondering “how slow is this?” IIUC, Kimi K3 on RTX 6000 Ada (48GB) takes 292 s/token https://github.com/lyogavin/airllm/releases/tag/v3.1.0
6.
▲
by
imenani
2mo ago
The relation to current RLVR methods I think is interesting, they do discuss it a bit but I would be curious to see more about this as well. Quote from the paper: Exploration beyond Pretraining. The mode collapse XMs address during pretrain
7.
▲
by
imenani
2mo ago
Nice presentation of the list! I'd recommend watching a few of his talks/podcasts before during reading these to get the overview and how all the bits in these works tie together. https://www.dwarkesh.com/p/il
8.
▲
by
imenani
4mo ago
https://lwn.net/Articles/1065620/
9.
▲
by
imenani
4mo ago
https://xcancel.com/tdietterich/status/2055000956144935055
10.
▲
by
imenani
4mo ago
The author discussed this here four days ago https://news.ycombinator.com/item?id=48077663
11.
▲
by
imenani
5mo ago
With the benefit of hindsight, perhaps much of this was Claude Mythos? The model was deployed internally since Feb
12.
▲
by
imenani
6mo ago
Agreed. LLMs have helped me achieve much deeper reading, _when directed to do so_. Asking an LLM to “Teach me Socratically about this paper/code. One question at a time”, usually allows me to get a much deeper reading of the material t
13.
▲
by
imenani
1y ago
Each of these models has a thinking/reasoning variant and a default non-thinking variant. I would expect the reasoning variants (o3 or “GPT5 Thinking”, Gemini DeepThink, Claude with Extended Thinking, etc) to do better at this. I thin
14.
▲
by
imenani
1y ago
As far as I can tell they don’t say which LLM they used which is kind of a shame as there is a huge range of capabilities even in newly released LLMs (e.g. reasoning vs not).
15.
▲
by
imenani
1y ago
They fix the temperature at T=0.6 for all k for all models, even though their own Figure 10 shows that RL model benefits from higher temperatures. I would buy the overall claim much more if they swept of temperature parameter for each k and