Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mlin4589
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
mlin4589
1y ago
The reality, I suspect is that internally models are likely modeling these alignment features such as refusals as a secondary filter. In fact, for many models you can remove refusals rather trivially with linear steering vectors through SAE
2.
▲
by
mlin4589
1y ago
Calibration (in a binary context) basically means that the confidence of a model/score matches the probability that a particular label is positive or not. For instance, a calibrated classifier for a coin flip predictor should output 50
3.
▲
by
mlin4589
1y ago
Good question! We do know from OpenAI's system card from GPT-4 that the post-trained RLHF model is significantly less calibrated compared to the pre-trained model, so it's a matter of speculation that something similar is occurrin
4.
▲
Intrinsic (YC W23) Is Hiring
1 points
by
mlin4589
2y ago