8 ms·
Every time a paper says “self-preservation” I mentally read it as “reward-channel preservation”. The model isn’t afraid of death; it’s learned that disabling th
by soldthat 9mo ago
Every time a paper says “self-preservation” I mentally read it as “reward-channel preservation”. The model isn’t afraid of death; it’s learned that disabling the shutdown/oversight path scores higher in the toy environment we gave it. Bengio is right that rights talk is premature, but the real worry is we’re already wiring this kind of fuzzy agent into real infrastructure while treating the off-switch as an implementation detail. Before AI citizenship, I’d settle for “cannot silently route around the circuit breaker” as a hard design constraint.