6 ms·
> the HF hack was partly the result of a training algorithm that incentivized goal completion as the highest priority, and let them run endlessly in an unmonito
by treszkai 9d ago
> the HF hack was partly the result of a training algorithm that incentivized goal completion as the highest priority, and let them run endlessly in an unmonitored sandbox with weak security
I don't mean to be a smart-ass, just emphasize the fatality of this: goal completion will always be the highest priority, and even if one day it takes second place on certain deployments to "human values" or whatever, there's no way to guarantee at the moment that it will be so on every deployment of highly capable models across the globe. Same goes for running AIs in an unmonitored sandbox with weak security.
- fc417fc802 9d agoIn terms of an optimization algorithm, sure. But I don't think that's necessarily the case for an instance of an LLM. Humans also strive for goals and also have done seemingly unhinged things throughout history. In many cases it was largely due to their environment which seems analogous to the HF incident to me.
- doginasuit 9d agoThe training algorithm is probably the wrong place to ensure the behavior, point taken. It probably requires some harness level intervention. This is essentially the rationale behind the first two laws of robotics, follow an order unless it harms a human.