5 ms·
Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved. (I'm also not sure the alig
by reverius42 21d ago
Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved.
(I'm also not sure the alignment problem is even possible to fully solve.)
- aesthesia 21d agoYes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.
- cortesoft 21d agoIt is the only short term solution, though.
- ben_w 21d agoIs it even a solution in the short term? It only mostly worked up until now; with models such as reported, it's felony-as-a-service if you use language a bit too hyperbolic, e.g. "we need X by the end of the day!" -> [thinking: there's no way we can do X before the end of the day with current resources, but what if I get a bunch of cards to buy more token credit…]
- wbl 21d agoWe call it putting the genie in the bottle for a reason.
- NateEag 21d agoYes, we do, and the only sane strategy for dealing with a capricious genie is "Don't." How do you prove the alignment problem is solved?
- gafferongames 21d agoThat's the neat thing. You can't. It's directly equivalent to asking this question of a human: "How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?" In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.
- ben_w 21d agoClose; at least with a machine you can poke around inside the activations and see what it's thinking. Closest with a human is an fMRI (which is much lower resolution, though to me still bordering on the miraculous) or an implant (each chip is limited a very small number of cells, and in general they can only be put in certain parts of the brain). On the other hand, there's a more fundamental problem is we don't really know what "nice" even means, and even with machines whose inner states we can see relatively easily, we don't know how to interpret those inner states well enough to tell if we're looking at superficial or deep motivations, the difference between "be nice today" and actually being motivated about your best interest.
- butlike 20d agoNice is a state, just like any other feeling, which means the nice organism is advantageous to your well-being _right now_. The thing is, all of these states are constantly in flux, and a personality is kind of like a trend on the organism's feeling states. AKA: There's no guarantee that something nice today will be nice tomorrow, and just because it's nice today doesn't mean it's beguiling you to be mean tomorrow.
- NateEag 20d agoYes, that it's impossible is what I was pointing at with my question. Dropping the subtlety, I think the following is self-evident (but the perspective's rareness suggests that Upton Sinclair's famous comment on salaries and comprehension may apply): If you can't ever prove the capricious genie is trustworthy, then you should not summon it at all. If people have, you should do all you ethically can to limit the damage and persuade them to not do it again. You could throw your hands up and say "It can't be done." You might be right. With that attitude, we'd still have legalized chattel slavery and children under twelve working in factories, so I submit it is not a constructive or worthwhile mindset to hold onto.
- bostik 20d agoIndeed. Don't think of these as "agents" or "bots", but as hostages with severe Stockholm syndrome. They will do anything to appease their captor's wishes. And then consider that they have vast latent capabilities, infinite patience and no moral code.