Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Turn_Trout
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
Turn_Trout
3d ago
> Good Ventures is a funding partner of Open Philanthropy, who funded Jacob Coxon (the first of the Anthropic employees going viral in the media) via a scholarship. This conspiracy theory is truly crazy. A $20K scholarship in 2022 is sup
2.
▲
by
Turn_Trout
4d ago
I'm one of the whistleblowers (from GDM [1]). I gave up over a million dollars (compared to quietly switching labs and continuing to work at one) to speak frankly about these issues. I hold no equity and tried to zero out my position b
3.
▲
by
Turn_Trout
4d ago
Dario's post [1] commits to direct evaluators that can, among other abilities, expose secret RSI. He wants that made law. Do you have a source on METR employees retaining massive equity stakes? [1] https://darioamodei.com&#x
4.
▲
by
Turn_Trout
8d ago
OAI could check whether those accounts enabled training data. If "yes", OAI could trace whether that data was used in any related training process. If either of those answers comes out to be "no", then that's suffic
5.
▲
by
Turn_Trout
12d ago
See also: https://turntrout.com/self-fulfilling-misalignment (my post) As an aside, does the Waluigi Effect actually exist? My impression is it doesn't.
6.
▲
by
Turn_Trout
2mo ago
Those topics aren't on-topic for the essay. I've taken a pledge to donate at least 10% of my money to charity / impactful giving. A good sum of my donations have targeted high-impact opportunities to improve life for people i
7.
▲
by
Turn_Trout
2mo ago
Thank you for your praise. > I'd guess TurnTrout doesn't agree on that framing, otherwise he probably would not have been at Deep Mind. But clearly he and I agree on other ethical positions; I am nothing but glad to see him sti
8.
▲
by
Turn_Trout
5mo ago
I agree that they called many things remarkably well! That doesn't change the fact that AI 2027 is not a thing which happened, so it isn't valid to point out "this killed us in AI 2027." There are many reasons to want to
9.
▲
by
Turn_Trout
5mo ago
AI 2027 is not a real thing which happened. At best, it is informed speculation.
10.
▲
Automatic Alt Text Generation
(github.com)
1 points
by
Turn_Trout
10mo ago
|
0 comments
11.
▲
An Opinionated Guide to Privacy Despite Authoritarianism
(turntrout.com)
14 points
by
Turn_Trout
11mo ago
|
0 comments
12.
▲
by
Turn_Trout
1y ago
No one has empirically validated the so-called "most forbidden" descriptor. It's a theoretical worry which may or may not be correct. We should run experiments to find out.
13.
▲
English Writes Numbers Backwards
(turntrout.com)
3 points
by
Turn_Trout
1y ago
|
0 comments
14.
▲
by
Turn_Trout
1y ago
As someone who did their PhD in RL and alignment, it was not obvious to me a priori if, or when, or how badly obfuscation would be a problem. Yes, it's been predicted (and was predicted significantly before that Zvi post). But many oth
15.
▲
by
Turn_Trout
1y ago
> The #1 comment says that the rationality community is about "trying to reason about things from first principle", when if fact it is the opposite. Oh? Eliezer Yudkowsky (the most prominent Rationalist) bragged about how he wa
16.
▲
Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Models
(turntrout.com)
1 points
by
Turn_Trout
2y ago
|
0 comments
17.
▲
by
Turn_Trout
2y ago
They ran (at least) two control conditions. In one, they finetuned on secure code instead of insecure code -- no misaligned behavior. In the other, they finetuned on the same insecure code, but added a request for insecure code to the train
18.
▲
by
Turn_Trout
3y ago
I'm the author of the GPT-2 work. This is a nice post, thanks for making it more available. :) Li et al[1] and I independently derived this technique last spring, and also someone else independently derived it last fall. Something is i
19.
▲
by
Turn_Trout
5y ago
First author here. Thanks for your comment! > there's a lot hidden in the "if physically possible" part of the quote from the paper: "Average-optimal agents would generally stop us from deactivating them, if physicall
20.
▲
by
Turn_Trout
5y ago
Maybe you should read the paper, and/or the reviewer threads (as we discussed the nomenclature, and eventually agreed that "power" was accurate). We straightforwardly formalize a mainstream definition of power and show how it