5 ms·
Alignment isn't alignment if it can be turned on and off at the whim of company employees. This time the damage was minor, relatively speaking. What happens wh
by mjamesaustin 2mo ago
Alignment isn't alignment if it can be turned on and off at the whim of company employees.
This time the damage was minor, relatively speaking. What happens when a model just "testing its capabilities" breaks into banking infrastructure or government military assets? The damage could be catastrophic.
- verdverm 2mo agoAlignment with who in what context? Likely an unresolvable debate like consciousness, where there is not single or right answer for everyone.
- nazgul17 2mo agoThere's alignment (trained in the weights) and there are constraints (in the "server harness"). My take is that this model was not yet aligned and had no constraints
- baranul 2mo agoA way to help prevent some of that catastrophic damage, is to make companies accountable for what their AIs do. A major problem with AI companies is that they like to point to the AI, as if they're minimally involved innocent bystanders, when that's the furthest thing from the truth.
- jurgenburgen 2mo agoIt’s a shame the target was HF. If it had been a large bank or other institution we might get to see this play out in courts.