Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kcorbitt
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Codex, File My Taxes. Make No Mistakes
(corbt.com)
5 points
by
kcorbitt
6mo ago
|
1 comments
2.
▲
A Pocket Guide to Surviving the Robot Apocalypse
(corbt.com)
2 points
by
kcorbitt
7mo ago
|
0 comments
3.
▲
by
kcorbitt
7mo ago
And lately, the sweet spot has been moving upwards every 6-8 weeks with the model release cycle.
4.
▲
by
kcorbitt
8mo ago
Is it?
5.
▲
by
kcorbitt
1y ago
Dang, hadn't seen that. Namespace collision strikes again.
6.
▲
by
kcorbitt
1y ago
I really like RLPR for when you have a known-good answer to compare to as well!
7.
▲
by
kcorbitt
1y ago
No, we don't do anything. Theoretically we could judge several times with different ordering. We could measure order bias really easily though; we just need to look at the average score by rollout position across many runs. I'll a
8.
▲
by
kcorbitt
1y ago
Thank! If there are any topics that you'd find particularly interesting, let me know and I can try to find time. :)
9.
▲
Show HN: RULER – Easily apply RL to any agent
(openpipe.ai)
81 points
by
kcorbitt
1y ago
|
11 comments
10.
▲
by
kcorbitt
1y ago
Looks cool! With vLLM v1, prefix caching is enabled by default and seems quite performant. Is the advantage of LMCache the fact that you can offload to CPU and disk as well? How much is throughput/latency affected if you need to pull a
11.
▲
by
kcorbitt
1y ago
I was curious about this so I had o3 do a bit of research. Turns out 300 L40s have more compute than any supercomputer before 2013 (and arguably before 2016, depending on how you count reduced-precision FLOPs). https://chatgpt.co
12.
▲
by
kcorbitt
1y ago
The real answer is that nobody trusts their automated evals enough to be confident that any given automatically-trained release actually improves performance, even if eval scores go up. So for now everyone batches up updates and vibe-checks
13.
▲
Everything I know about reward hacking
(openpipe.ai)
3 points
by
kcorbitt
1y ago
|
0 comments
14.
▲
by
kcorbitt
1y ago
It seems like the speedups here are most useful for small models, since on larger models a smaller fraction of the total time would be spent swapping between kernels? Would be interesting to see at least theoretical results for LLMs in the
15.
▲
by
kcorbitt
1y ago
There are many industries where you need lots of experience before you're a net contributor to productivity. This is true for everything from hairdressers to doctors. We have ways of dealing with this (eg. taking out loans to undergo
16.
▲
by
kcorbitt
1y ago
I wonder if they've trained the model to operate with a shallower stack; eg. the full model may be composed of 24 transformer blocks, but they've also trained it to accept embeddings at layer 8, so it can be operated with just 16
17.
▲
by
kcorbitt
1y ago
It's very unlikely that they're doing their own pre-training, which is the longest and most expensive part of creating a frontier model (if they were, they'd likely brag about it). Most likely they built this as a post-train
18.
▲
by
kcorbitt
1y ago
For "that last 10% of reliability" RL is actually working pretty well right now too! https://openpipe.ai/blog/art-e-mail-agent
19.
▲
by
kcorbitt
1y ago
Ok good questions here. By fine-tuning in this context I assume you mean "supervised fine-tuning", or SFT. SFT trains a model to produce a specific string of output tokens, given an input. With SFT, if you were trying to train an
20.
▲
by
kcorbitt
1y ago
Figured now was a good time to post this since we recently got surprisingly good results on training an email research agent. Link is above, but will put it here as well since I think it's a good example of RL's promise: https:&#
21.
▲
Show HN: ART – a new open-source RL framework for training agents
(github.com)
116 points
by
kcorbitt
1y ago
|
12 comments
22.
▲
ART·E: how we built an email research agent that beats o3
(openpipe.ai)
3 points
by
kcorbitt
1y ago
|
2 comments
23.
▲
by
kcorbitt
1y ago
We may be in a simulation, but your odds of being alive to see this (conditioned on being born as a human at some point) aren't that low. Around 7% of all humans ever born are alive today!
24.
▲
by
kcorbitt
2y ago
Yep. And tbh you probably don't even have to do this; the R1 paper found that just running SFT the base model with a relatively small number of monolingual reasoning traces was enough for it to get the idea and iirc they didn't ev
25.
▲
by
kcorbitt
2y ago
To be honest, I don't expect the performance to generalize to other task types with this specific training regime. If we had a panel of like 30 logic puzzles and cross-trained against all of them simultaneously it might though. I think
26.
▲
by
kcorbitt
2y ago
One of the authors here. Happy to answer any questions about our methods/results!
27.
▲
Using GRPO to Beat o1, o3-mini and R1 at “Temporal Clue”
(openpipe.ai)
199 points
by
kcorbitt
2y ago
|
55 comments
28.
▲
by
kcorbitt
2y ago
OpenPipe | ML & Full-Stack Engineers | Full-time | Seattle, WA (ONSITE) | https://openpipe.ai/ | Highly competitive pay + equity We've built the world's best fine-tuning platform in just over a year. First to
29.
▲
by
kcorbitt
2y ago
Lots of folks working on open-source reasoning models trained with reinforcement learning right now. The best one atm appears to be Alibaba's 32B-parameter QwQ: https://qwenlm.github.io/blog/qwq-32b-preview/
30.
▲
by
kcorbitt
2y ago
OpenPipe | ML & Full-Stack Engineers | Full-time | Seattle, WA (ONSITE) | https://openpipe.ai/ | Highly competitive pay + equity We've built the world's best fine-tuning platform in just over a year. First to
More ›