Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ag8
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
ag8
5mo ago
I find this paragraph to be odd: "Wavelengths as low as 13.5 nanometers can achieve more precise patterns in a single exposure. In fact, extreme ultraviolet lithography can combine three or four photolithography patterning cycles into
2.
▲
by
ag8
7mo ago
You're right; I should've been more precise. However, we have tools for dealing with this—that's what quality-adjusted life-years are for! I don't contest that surgeries often significantly increase QALYs, and may do so
3.
▲
by
ag8
7mo ago
Lol, I just care a lot about saving as many lives as I can; the most effective charities I've been able to find good evidence on save one life for $6–8k. If Watsi had a credible claim at being able to save lives 10x cheaper I would red
4.
▲
by
ag8
7mo ago
Watsi seems to be doing great work, but the title—"you helped save 33k lives"—reads as misleading to me. I guess "helped" could be doing a lot of heavy lifting here, but I would be incredibly surprised if the counterfact
5.
▲
Gourmand Syndrome
(en.wikipedia.org)
27 points
by
ag8
8mo ago
|
9 comments
6.
▲
by
ag8
8mo ago
https://andrew.gr
7.
▲
guys why does armenian completely break Claude
(twitter.com)
99 points
by
ag8
8mo ago
|
65 comments
8.
▲
Sampling at negative temperature
(cavendishlabs.org)
203 points
by
ag8
8mo ago
|
60 comments
9.
▲
Perfectly Replicating Coca Cola [video]
(youtube.com)
1 points
by
ag8
8mo ago
|
1 comments
10.
▲
by
ag8
9mo ago
Not 13?
11.
▲
by
ag8
9mo ago
This is a cool setup, but naively it feels like it would require hundreds of thousands of hours of data to train a decent generalizable model that would be useful for consumers. Are there plans to scale this up, or is there reason to believ
12.
▲
Po.ta.to
(po.ta.to)
4 points
by
ag8
10mo ago
|
2 comments
13.
▲
Scaling pretraining affects RL sample efficiency
(runrl.com)
1 points
by
ag8
11mo ago
|
0 comments
14.
▲
Systematically generating tests that would have caught Anthropic's top‑K bug
(theorem.dev)
2 points
by
ag8
11mo ago
|
0 comments
15.
▲
by
ag8
1y ago
Yeah, not sure why the HN backend changed it...
16.
▲
Tinker
(2b4fdb18.connectionism.pages.dev)
4 points
by
ag8
1y ago
|
2 comments
17.
▲
Training Qwen to answer briefly yet intelligently using feedback control
(runrl.com)
4 points
by
ag8
1y ago
|
0 comments
18.
▲
by
ag8
1y ago
A) You could have an additional field in the jsonl file which says which rubric to use; then, your reward function could access this via `kwargs["rubric"]` and return a reward based on that example's preferred rubric; B) curr
19.
▲
by
ag8
1y ago
Having an RL agent that's really good at search across some space sounds very powerful in general; "proofs-as-search" make this an appealing target. Back in the day, when I did more fundamental RL research, we worked on an ex
20.
▲
by
ag8
1y ago
we should publish some; the high-order effect seems to be that LoRAs significantly hurt small model performance vs FFT, with less of an effect for large models. This is maybe because large models have more built-in skills and thus a LoRA su
21.
▲
by
ag8
1y ago
Thanks! Our goal is to make rl "just work" with completely automated GPU provisioning/algorithm selection/SFT-warm up, but giving people the ability to switch away from the defaults if they want to. The way tools current
22.
▲
by
ag8
1y ago
Yeah, for better or worse, the way the median startup interfaces with AI these days is through an LLM API, and that's what all the workflows are built around, so that's what we're targeting. Though, depending on what you'
23.
▲
by
ag8
1y ago
It's for any task that has an "eval", which is often verifiable tasks or ones that can be judged by LLMs (e.g. see [0]). There's also been recent work such as BRPO [1] and similar approaches to make more and more "n
24.
▲
by
ag8
1y ago
prompt optimization is very cool, and we use it for certain problems! The main goal with this launch is to democratize access to "the real thing"; in many cases, full RL allows you to get the last few percent in reliability for th
25.
▲
Launch HN: RunRL (YC X25) – Reinforcement learning as a service
(runrl.com)
71 points
by
ag8
1y ago
|
22 comments
26.
▲
Generating the Funniest Joke with RL
(runrl.com)
1 points
by
ag8
1y ago
|
0 comments
27.
▲
by
ag8
1y ago
wow, I used to make so many games with image maps back when I first learned HTML. One still survives: https://andrew.fi/beowulf/game/
28.
▲
by
ag8
1y ago
Reminds me of https://szge.ca , which comes with fake keyboard and fan noises:)
29.
▲
Gravity Chess
(gravity-chess.andrew.gr)
2 points
by
ag8
2y ago
|
1 comments
30.
▲
by
ag8
2y ago
Yeah, I'm a bit surprised that this has so many upvotes without any easily accessible text
More ›