Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
daveyoung
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
daveyoung
22d ago
do you have reference to where it was said?
2.
▲
by
daveyoung
22d ago
likely a distilled glm 5.3 that will punch within 20% of that at 2-3x less size. you'll find that capability is typically very jagged on models that are distilled
3.
▲
by
daveyoung
22d ago
Two potentials from my pov: 1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/