5 ms·
> So going to find the Vulnerability's description on a third party website is clear cut reward hacking that depends on what the prompt was, maybe they worded
by Nathanba 2mo ago
> So going to find the Vulnerability's description on a third party website is clear cut reward hacking
that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.
- pixl97 2mo agoI don't quite understand how that changes anything? In the story of the paperclip maximizer it boils down to >But for all its sophistication, it understood only the simple objective that had been programmed into it: it must at all costs maximize the number of paperclips.
- dinfinity 2mo ago1. They explicitly disabled the "don't be evil" protections: "We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity." 2. Hacking HuggingFace to get to its datasets is a far cry from "consume/kill all humans". It's very very specific to the task at hand and easily predicted given the lack of guardrails.
- dinfinity 2mo ago> that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest. That is an interesting question. If the prompt included "Do not break out of the sandbox we've provided you. Do not use information retrieved from outside the sandbox. All answers that were provided in this manner are invalid and will score 0 points.", would this still have happened?
- HeatrayEnjoyer 2mo agoSo...? We should not construct a machine that is one bad prompt away from causing catastrophe.