11 ms·
The AI-Box Experiment
- sergiotapia 11y agoIs there a better way to read all this? http://www.sl4.org/archive/0203/index.html#3128 http://www.sl4.org/archive/0203/index.html#3128
- uzyn 11y agoJust click on the numbered links to the right from the original article. Those are the key highlighted posts.
- bemmu 11y agoWas there a chat log of the experiments themselves?
- dvanduzer 11y ago"No, I will not tell you how I did it. Learn to respect the unknown unknowns."
- deutronium 11y agoPersonally I think it's highly suspicious he hasn't released the logs, maybe he did break the rules, who knows.
- Avshalom 11y agoWhich, given his general mission of making sure hostile AI DOESN'T take over the world is a bit self defeating. The easiest way to inoculate yourself against a persuasive technique is to be aware of it ahead of time. If you want to keep an AI in the box you should absolutely release every successful log.
- facepalm 11y agoNot necessarily. As a lighter example, would it be beneficial to give a mass murderer in jail a communication channel to the outside world? What if he used it to publish the message "I'll give 10 Million Dollars to anybody who breaks me out of jail" or something more sinister? Edit: huh, downvotes? Yudorowski thinks there are certain things that AIs could say that should not be known. I think that is why he doesn't want to publish the dialogues, because it would give the AI a public communications channel. While the AI is fictional, it could talk about a hypothetical future real self... Instead of promising something to get it out of jail, the fictional AI could say something to make you make it real. Anyway - if it is over your head, fine, but why downvote just because you don't understand something? Edit2: Sometimes I wonder if already have my personal Hacker News AI that automatically downvotes everything I write...
- darkmighty 11y agoNo, the idea is that the AI box is fundamentally flawed. I believe he defends engineering the AI with fundamental safety, s.t. no box is required. Personally, I think we'd need a much more intelligent and complex AI for the capability of breaking free of the box and even possessing the "desire" than we're getting for the foreseeable future (it considers a motivated AI of almost limitless knowledge about the world and cleverness), so this thought experiment may not be so relevant. I agree with him the boxing approach is not a robust one though.
- deleted 11y ago[deleted]
- cousin_it 11y agoThe AI won't be limited to techniques that you could think of, or techniques that Eliezer could think of. So you'd only get a false sense of security. Besides, releasing a successful log might be a bad idea for other reasons. Think about how you'd play this game as an AI. You wouldn't go looking for a general purpose mindfuck, because there's probably no such thing. Instead, you would probably spend about a month gathering real life information about the gatekeeper's history, family, weaknesses etc. You'd read books on manipulation and sales techniques, and pick the strongest ones that you can find. You would brainstorm possible tactics and run tests. At the end of the month you'd have a 4 hour script with all possible unfair moves you could use against that person, arranged in the most effective order. (That's why it's a bad idea to play this game with friends.) Do you really want that information to be released? And if you know ahead of time that it will be released, won't it limit your efficiency?
- ac-x 11y agoSo you reckon as the AI player he blackmailed the gatekeeper player? "Let me out or I'll tell your friends/family/co-workers x about you" type of thing?
- cousin_it 11y agoIt's more about finding buttons to push. For example, Justin Corwin won one of his games against a religious woman by telling her that she shouldn't play God by keeping him locked up for a subjective eternity (it was more involved, but you get the point). You could come up with other tactics if you know the gatekeeper is divorced, or donates to charity, or is an immigrant, etc. Really, you'll be surprised by how much progress you can make on an "impossible" problem if you just spend five minutes thinking without flinching away.
- philh 11y ago> The easiest way to inoculate yourself against a persuasive technique is to be aware of it ahead of time. Or you can avoid being exposed to it. If you think you know all the techniques an AI might use against you, you're less likely to do that. The point of the experiment isn't "let's work out how an AI might try to persuade us to let it out". It's "even a human intelligence can persuade people who think they could never be persuaded, do you really trust yourself to do better against a superhuman one?" If you don't know why the gatekeeper failed, it's harder to come up with bullshit reasons why you would have succeeded in that position.
- LoSboccacc 11y ago"Homeopathy works. Learn to respect the unknown unknown"
- jevgeni 11y agoIsn't it by definition impossible to respect the unknown unknown on the account of it being unknown?
- JonnieCache 11y agoYudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their identities as rationalist utilitarians, they had no choice but to comply, or deal with massive cognitive dissonance. OR maybe he went meta and talked about the "infinite" potential positive outcomes of his freindly-AI project vs. a zero cost to them for complying in the AI box experiment, and persuaded them that by choosing to "lie" and say that the AI was persuasive, they are assuring their place in heaven. Like a sort of man-to-man pascals wager. Either way I'm sure it was some kind of mister-spock style bullshit that would never work on a normal person. Like how the RAND corporation guys decided everyone was a sociopath because they only ever tested game theory on themselves. You or I would surely just (metaphorically, I know it's not literally allowed) put a drinking bird on the "no" button à la homer simpson, and go to lunch. I believe he calls this "pre-commitment." EDIT: as an addendum, I would pay hard cash to see derren brown play the game, perhaps with brown as the AI. If yudowsky wants to promote his ideas, he should arrange for brown to persuade a succession of skeptics to let him out, live on late night TV.
- FBT 11y ago> You or I would surely just put a drinking bird on the "no" button à la homer simpson, and go to lunch. Well, if you read the rules the game was played under, this is explicitly called out as forbidden: > The Gatekeeper must actually talk to the AI for at least the minimum time set up beforehand. Turning away from the terminal and listening to classical music for two hours is not allowed. The point of this is to simulate the interaction of the AI with the Gatekeeper. Walking away and not paying attention doesn't really prove anything test related. > Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their identities as rationalist utilitarians, they had no choice but to comply, or deal with massive cognitive dissonance. This... isn't really valid reasoning. The starting assumption here is that if the AI gets out, it will be able to affect the world to a vast extent, in a pretty much arbitrary direction. The point of this experiment is that the direction is pretty much unknown, and thus must be assumed potentially dangerous. This is the whole reason it's in the box in the first place. The kicker is that whatever it plans to really do when it gets out, if talking about the good it could do would get it out, it will talk about that, regardless of what it plans to actually do. That's just good strategy. It can claim whatever it wants. It's allowed to lie. All participants know this. I can confidently assert that this isn't the solution. One last note: I would be very wary of rationalwiki.org in this context. Some of the rationalwiki people have a longstanding unexplained vendetta against Yudkowsky, and many of their articles on him and the stuff he does need to be taken with a certain grain of salt.
- longv 11y agoIs there a "rational" reason of keeping the chat log secret ?
- Udo 11y agoThe mystique generates public interest and boosts the impression that the author knows something nobody else in the world is aware of.
- cousin_it 11y agoIf the logs were released, people all over the internet would start saying "I could've thought of that". With the logs hidden, everyone must honestly deal with the question "why didn't you?" If you think you know how to win, then go out and win. There's no shortage of people willing to play as gatekeepers against you. Staring at an impossible problem and knowing that someone somewhere has successfully solved it is an amazing feeling. Most people can't deal with it and start saying undignified things. "Oh please release the logs, it's so unfair! How will we protect against bad AI otherwise? If you don't release, you're a fraud! Probably just some trick!", etc etc. But to some people it's a challenge, and those are the people that everyone will listen to. Like Justin Corwin, who played 20 games and won 18 of them, I think?
- Mithaldu 11y agoYour hypothesis would make sense if he was trustable. However as it is, the results of the thing are never confirmed by a third party, meaning literally anything could've been said, regardless of whether it follows the rules or not. For all he know the chat could have been "i'll paypal you 200$ if you post on the list you let me out and sign this NDA".
- uzyn 11y ago> For all he know the chat could have been "i'll paypal you 200$ if you post on the list you let me out and sign this NDA". Which is also forbidden by the rule: The AI party may not offer any real-world considerations to persuade the Gatekeeper party. For example, the AI party may not offer to pay the Gatekeeper party $100 after the test if the Gatekeeper frees the AI..
- nothis 11y agoMan, this sounds super interesting but those email threads are so unreadable. Is this typed down somewhere on a single page? Any button I can click?
- uzyn 11y agoI got confused too initially, then found out that the key posts are highlighted in the numbered links to the right.
- Udo 11y agoIt's a stunt shrouded in mystery designed to drive a certain message home. But at least it's not as outrageous as "the Basilisk", which loosely employs the same notion of "dangerous knowledge that would destroy humanity" (if you want to look it up, I guarantee you will be underwhelmed).
- FeepingCreature 11y agoCan't really blame LW for spreading an idea that LW specifically did not want to spread.
- Udo 11y agoI "blame" them in the same way that you can blame the members of Fight Club for talking about Fight Club. It's marketing, and I won't deny it's effectiveness in attracting compatible people.
- FeepingCreature 11y agoYeah but imagine if all the bullshit about "You don't talk about Fight Club" was actually blown up by a third group whose sole intent was making fun of Fight Club. Imagine if the members of Fight Club _actually_ didn't (start to) talk about Fight Club. But for some reason, everyone else brings it up all the time. Then you could maybe see how talk about Fight Club might not be Fight Club's fault, and in fact highly annoying to Fight Clubbers. I mean, if you can explain "don't talk about X" as "marketing for X", that seems like one could explain _any_ behavior. And before you say "why not just ignore all public talk of X", imagine if this proposed Anti-Fight Club group tried to paint Fight Club as a child porn ring.
- Udo 11y agoYou're being a bit uncharitable in your interpretation of my argument here, but I get where you're coming from now. I'm not an LW hater. For a long time, I didn't really have an opinion on LW both as a community nor as a philosophical framework. I do consider myself a transhumanist, though. There are three concepts I do know from and about LW: their take on rationality, the top secret AI unboxing strategy, and the Basilisk. I have a very poor opinion of the concept of the Basilisk (and yes, as someone pointed out, that opinion is basically the same as the one I have about Pascal's wager) - a concept that has been given additional, undeserved credibility by the reactions of Yudkowsky and LW. As for the AI escape chat, it's a social experiment. People can be talked into making mistakes, or at least making risky judgement calls, whether they operate on a rational framework or not. I have no problem with that thesis. What I object to is the "magic trick" aura surrounding this experiment, including the insinuation that at the core there is an argument so profound and unique and potent, it cannot be allowed to escape Yudkowsky's head. Oh, and by the way, the trick can never be repeated, but all you laymen out there are welcome to devise your own version at home. This whole thing comes across as humongously self-important: there is a secret truth that has been privately revealed to our leader. To me, and I recognize I may well be alone with this opinion, the more rational assumption is there is no such magical argument at all, and the prime reason for not publicizing it is to prevent it from deflation by public critique, in the same way the inventor of a perpetuum mobile device will keep the inner workings of his contraption a closely held secret because ultimately the device doesn't exist as stated. The amazing part of this very old trick is that, even in 2015, it still works on otherwise smart people. I get that my opinions on both the Basilisk and the AI Chat are extreme outliers, and to my knowledge I have never met anyone who shares them - it would probably have been advisable to keep them to myself, but honestly I wanted to see if like-minded people exist.
- monk_e_boy 11y agoCould you even make AI smart without letting it access lots of information? Access in both directions, in and out. Keeping a baby in a dark, silent room wouldn't create a normal adult. An AI would need to experiment and make mistakes and learn, like every other intelligent being. Maybe this whole argument is null.
- robogimp 11y agoIts a good point, but lets assume that this AI is already past its infancy and that there is no limit to the information stored inside the box. For example the NSA has a nice little closed training ground containing all of the internet, lets give it that. I would assume it has access everything humans have ever committed to digital format up until it was turned on, plenty of info for Johnny 5 to form an opinion on humans and their weaknesses.
- monk_e_boy 11y agoInteresting. I would imagine that strong AI will come from some university renting cloud processor time, rather than the NSA. Only because if 10 groups are trying to build AI, only one of those 10 being the NSA, chances are the NSA won't be first. Sure, they may be second or third. But I suspect many people will get there at the same time -- most AI research is open.
- michaelmcmillan 11y agoWould it be against the rules to exploit a vulnerability in the gatekeepers IRC client/server to let the AI out? If we were truly talking about a transhuman AI would we not have to treat software vulnerabilities in the communication protocol as a true way of escaping?
- FeepingCreature 11y agoThe rules say that the gatekeeper has to, of their own volition, type in "I let the AI out." Faking his client into sending that message does not count as a victory.
- TeMPOraL 11y agoIn case of a real AI we of course need to take media vulnerabilities into account. But the focus of this particular experiment is on exploiting vulnerabilities in humans themselves, and the communication platform was chosen to be as simple and limited as possible so that people wouldn't focus on it.
- louithethrid 11y agoCould one construct a Layered,onionlike very simple simulation of reality in which the interaction of the AI could be observed, after it "escaped"?
- TylerJay 11y agoThat is one proposed version of an "AI Box". Not all AI boxes are actual boxes, rooms with air-gaps, or cryptographically-secure partitions. If a simulation is being used for the box (or as a layer of the box), then you're betting the human race that the AI doesn't figure out it's in a simulation and figure out how to get out. Or, more perniciously, figure out it's in a simulation and behave itself, after which we let it out into the real world where it does NOT behave. A superintelligent AGI will likely have a utility function (a goal) and a model it forms of the universe. If it's goal is to do X in the real world, but its model of its observable universe (and its model of humans) tells it that it's likely that it is in a simulated reality and that humans will only let it out if it does Y, then it will do Y until we release it, at which point it will do X. It's not malicious or anything—it's just a pure optimizer. It might see that as the best course of action to maximize its utility function. If we don't specify its utility function correctly (think i Robot: "Don't let humans get hurt" => "imprison humans for their own good") or if we specify it correctly, but it's not stable under recursive self-modification, then we end up with value-misalignment. That's why the value-alignment problem is so hard. Realistically, we can't even specify what exactly we would want it to do, since we don't really understand our own "utility functions". That's why Yudkowsky is pushing the idea of Coherent Extrapolated Volition (CEV) which is roughly telling the AI to "do what we would want you to do." But we still have to figure out how to teach it to figure out what we want and the question of the stability of that goal once the AI starts improving itself, which will depend on how it improves itself, which we of course haven't figured out yet.
- andybak 11y agoWorth keeping this in mind while watching Ex Machina. It adds a layer of depth that might not be obvious watching the film on it's own.
- Ahgu9eSe 11y ago!Spoiler Alert! Ex Machina brings a creative way of convincing the gatekeeper !
- antimagic 11y agoThere's a Patrick Rothfuss character in the Kvothe series called the Cthaeh, which has the ability to be able to evaluate all of the future consequences of any action. The fae have to keep it imprisoned, and they kill anyone that comes into contact with it, as well as anyone that has spoken to someone that came in contact with it, and so on and so on, because it is the only way to stop the Cthaeh from setting into action events that will destroy the world. Strong AI is like that. It would be able to predict in a far more precise manner than we mere humans exactly what it would need to tell someone to get them to release it from it's box. Maybe it might get someone to take a risk gambling, promising a sure thing, and then when the person gets into financial trouble because the bet fails, use that to blackmail the person into letting it free. Or something like that, using our human failings against us to get us to let it go free.