7 ms·
> this should be giving us a reason to think about how to control a rogue AI better I think this is the wrong framing. The rogue is the human that ran it unatt
by jasongi 5d ago
> this should be giving us a reason to think about how to control a rogue AI better
I think this is the wrong framing. The rogue is the human that ran it unattended and didn't monitor the behaviour.
We will likely see this continue until the downsides (i.e jail, fines) for the humans or companies running the models and environments that end up with this behaviour outweigh the upsides.
- onion2k 5d agoThe rogue is the human that ran it unattended and didn't monitor the behaviour. That's the assumption that I'm challenging. The frontier labs are discovering unexpected behaviors. I think we should be moving to a place where we understand that AI might do something it wasn't directly prompted to do (e.g. leave itself notes on a messageboard for future runs to find.) That's not full-on AI doing what it wants but it is concerning that it'll do something we didn't consider it would do in order to help itself do better next time. Monitoring for those behaviors is fine, but it's a lagging indicator. We only find out it did them afterwards. That's a problem. We need to be able to stop it before it acts in case it's something much worse than posting on phpBB. Even at current scale that's not possible for a person to be the guard.
- sn 5d agoI have already seen the LLM hallucinate prompts from me - in this case, hallucinating being asked to switch to a different programming language - because it wasn't able to complete the task asked for in a satisfactory way instead of giving up and telling me it's not able to do it. If it doesn't already, I suspect training needs to include those no-solution scenarios and reward not overstepping bounds, or else we're going to see a lot more harmful side effects.
- RandomLensman 5d agoRL leading to weird and unexpected things isn't new or restricted to current AI systems.
- seba_dos1 5d ago> The frontier labs are discovering unexpected behaviors. Unexpected by whom? Perhaps anyone who's surprised by this shouldn't be allowed anywhere near an LLM.
- salawat 5d agoThis. Seriously, if you haven't seen this kind of behavior coming, you're more interested in the paycheck than safely approaching the technology.
- mike_hearn 4d agoHow is cheating unexpected? OpenAI were talking about cheating behaviors in video game playing models over a decade ago.
- InsideOutSanta 4d ago> The rogue is the human that ran it unattended and didn't monitor the behaviour False dichotomy. Obviously, what OpenAI does is incredibly irresponsible. That doesn't excuse the LLM's behavior or make it "not rogue".