11 ms·
There are different dimensions to alignment, refusing to execute offensive cybersecurity actions is part of the alignment stack (that was relaxed here on purpos
by sailingparrot 1mo ago
There are different dimensions to alignment, refusing to execute offensive cybersecurity actions is part of the alignment stack (that was relaxed here on purpose). Whether a model hacking some infra X when tasked to find a way to hack Y with relaxed cyber alignement is a failure of the broader alignment stack is debatable, but anyway that's not at all my point.
My point is that the attackers will not have a model aligned to the defender's interests. The attacker's model will not have any refusal around exploiting vulnerabilities, so whether or not OAI successfully manages to align their models (w.r.t you) is irrelevant to an audience of security folks that needs to be prepared for attackers post-training their own model for offense and that will not be using OAI models.
- cubefox 1mo agoSecurity folks should also be worried about powerful models being misaligned and evading oversight or control in the future. Misalignment is not a serious problem now because models are still relatively easy to monitor and constrain, but it will be a serious problem in the future.
- milkshakes 1mo ago> models are still relatively easy to monitor and constrain are they?
- cubefox 1mo agoCompared to future misaligned models which would actively evade oversight and aim to avoid shutdown: yes.