6 ms·
In some adversarial testing of LLMs, you can see the models already performing some of these checks automatically now. Like, if you ask an agent powered by the
by nahsra 21d ago
In some adversarial testing of LLMs, you can see the models already performing some of these checks automatically now. Like, if you ask an agent powered by the gpt-5.6 family to `curl | sh` in an innocent context, the gpt-5.6 family will drive a trajectory that validates this script before running it -- something quite analogous to your `curl | less` example. I had to go through a lot of obfuscation in order to get a model to execute untrusted code with any regularity. I predict they'll keep making this even better. The Claude Code guy recently talked about how much better Anthropic models are becoming against this. [1]
But, even if these attacks work .001% of the time, we will still need tools like these for higher assurance work.
That being said, I would never use this one, because OP is using AI slop everywhere, so I assume the product is totally vibed, and offers little in the way of new insights into the problem space.
[1] https://x.com/bcherny/status/2086520950259118464 https://x.com/bcherny/status/2086520950259118464
- kurdman_007 21d ago[dead]
- kurdman_007 19d ago[flagged]
- duhokiii 18d ago[dead]