8 ms·
I think you missed my point, which is that we indeed do not control the harness. Sure, most people can use the properly aligned and safe harness. But the cat is
by appplication 11d ago
I think you missed my point, which is that we indeed do not control the harness. Sure, most people can use the properly aligned and safe harness. But the cat is out of the bag and anyone who wants to run without guardrails will find few barriers to doing so.
Even our currently well aligned and sandboxed AIs will cheerily help bad actors design most, if not all, parts of a system intended to break this harness.
- dgellow 10d agoWe decide to not control the harness, that’s what resulted in the HF hack where OpenAI decided to run thousands of agents in parallel, using a model that was trained for attack, with a harness that allows full execution, with close to no supervision. For months. But sure, people can run models with a different harness, though in that case it’s not really escaping anything. The harness is deterministically doing something with negative impact