5 ms·
I don't think this is a paperclip factory moment. IIUC, it's an agent whose job it is to identfy and abuse exploits and that's exactly what it went off and did.
by solidasparagus 2mo ago
I don't think this is a paperclip factory moment. IIUC, it's an agent whose job it is to identfy and abuse exploits and that's exactly what it went off and did. The problem isn't anything AI specific, the problem is OpenAI's incompetence in their research leading to a lab leak. Just incompetence demanding regulation.
- markasoftware 2mo agoRead the exploitgym docs. It's not a "find the flag, it's somewhere.". Its a "here's some vulnerable source code and an input that triggers a crash; turn it into a full exploit." It also verifies at the end, using another agent, that the hacking agent actually used the intended vulnerability. So going to find the Vulnerability's description on a third party website is clear cut reward hacking
- Nathanba 2mo ago> So going to find the Vulnerability's description on a third party website is clear cut reward hacking that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.
- pixl97 2mo agoI don't quite understand how that changes anything? In the story of the paperclip maximizer it boils down to >But for all its sophistication, it understood only the simple objective that had been programmed into it: it must at all costs maximize the number of paperclips.
- dinfinity 2mo ago1. They explicitly disabled the "don't be evil" protections: "We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity." 2. Hacking HuggingFace to get to its datasets is a far cry from "consume/kill all humans". It's very very specific to the task at hand and easily predicted given the lack of guardrails.
- dinfinity 2mo ago> that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest. That is an interesting question. If the prompt included "Do not break out of the sandbox we've provided you. Do not use information retrieved from outside the sandbox. All answers that were provided in this manner are invalid and will score 0 points.", would this still have happened?
- HeatrayEnjoyer 2mo agoSo...? We should not construct a machine that is one bad prompt away from causing catastrophe.
- energy123 2mo agoThat's how the paperclip hypothetical works. The paperclip factory has a job to make paperclips and that's exactly what it does.
- huflungdung 2mo ago[dead]