6 ms·
Am I the only person reading the statistics in this announcement from Anthropic and the associated blog commentary and trying to work out how they possibly coul
by Silhouette 1mo ago
Am I the only person reading the statistics in this announcement from Anthropic and the associated blog commentary and trying to work out how they possibly couldn't imply that a significant number of dangerous commands are likely to be attempted every day these tools are in use and neither manual human review nor the auto classifier provided by Claude is anywhere near reliable in preventing them?
A lot of the discussion about these long sessions where agents are left to operate autonomously feels like listening to the increasingly drunk guy at the bar who says "I ran IT at that Fortune 100 place for a decade and we never had a single problem using a short but loose rule set for the firewall until last week someone destroyed our entire business in 27 minutes".
- cillian64 1mo agoI think by "dangerous" they include things like "makes an edit to a config file outside the current project", not just "wipes the production database". So "dangerous" commands just means things we should ask the user for confirmation, not commands which definitely cause irreversible damage.
- Silhouette 1mo agoEven if that's what they mean it would still be disturbing if commands that should have explicit user confirmation were being waved through - whether by users or the auto mode - and being run when they shouldn't. The statistics appear to suggest that this is not only possible but actually quite likely given the number of commands an agent running all day might propose to run. But if that failure mode is a possibility at all then IMHO the whole system is playing with fire.