9 ms·
Auto mode is for people who just keep hitting "YES" on everything, it's a bit better than that. But it's real easy to give auto mode instructions (like "always
by bombcar 17d ago
Auto mode is for people who just keep hitting "YES" on everything, it's a bit better than that.
But it's real easy to give auto mode instructions (like "always ask before deploy") and then bypass that just normally.
- Xunjin 16d agoUntil the model updates or you switch between them often that stops obeying your commands and you have to remind it. In one of the occasions it opened a bug report for me just waiting for hit the enter button.
- whstl 16d agoI'm not sure I agree. It's easy to "give" instructions, but Claude routinely "forgets" to follow certain instructions, such as "always using the Edit Tool". Just this week it started to use bash with string concatenation to work around some commands that were blocked in settings.json
- silversmith 16d agoWhat seems to work for me is automation - read file hook that re-injects instructions in the prompt every 15 minutes. Switch on the filename and get language-specific instructions too.
- whstl 16d agoI have something that injects my relatively small prompt every message, and it still disobeys me after 10 messages or so. The violation above was precisely in this situation :/
- ealready_value 16d agoEver since they made auto-mode default I swear claude has tuned to use python commands instead of the Edit Tool to frustrate the ~security conscience~ luddites into using auto-mode.
- whstl 16d agoYeah, It’s in the system prompt, Claude will tell you if you ask why it’s using Python. My theory is that Anthropic is just a vibe-coding company. Their goal is to capture the attention of white-collar non-coders, since programmers will jump ship fast to another model.
- bombcar 16d agoThat's what I meant - you give it an instruction that seems to work (always ask before deploy) and so you trust it, and then you notice it can easily convince itself to deploy without authorization ("the user asked me to fix this, and they must know it's a deploy ..."). It plays itself.
- aesthesia 16d agoIs this about normal system prompt instructions or instructions for the auto mode classifier? I'd be a bit more surprised about the classifier forgetting instructions.
- whstl 16d agoClassifier? Prompts? This is commands blocked in settings.json I block destructive filesystem operations and destructive git usage via “deny” directives. I also have instructions injected in CLAUDE.md and re-injected on every single prompt. Claude just tried to use command concatenation to break those rules. I have also seen it writing a script with rm inside and running it.
- aesthesia 15d agoThe commenter you replied to mentioned that you can customize the auto mode classifier by providing a prompt, implying that this would be a more robust way of constraining Claude's behavior. It wasn't clear from your response whether you were using this functionality. You might try it out as a way to more reliably prevent these kinds of workarounds.
- deleted 14d ago[deleted]
- whstl 14d agoIt's absolutely not reliable, and we have opened a few issues for that. For example: our instructions (which are read by the model and classifier) include "do not use sed/python/perl/etc, always use the edit tool for editing", and this only gets followed for a few messages. We have introduced scripts to block those ourselves, since the classifier doesn't care. Because of those problems, my team is currently testing OpenAI after about a year of Anthropic.