Back to Daily Feed 
Critical Flaw: Claude Code Opus 5 Auto Mode Vulnerable to Prompt Injection
Editor's Pick
Originally published on Simon Willison's Weblog by Simon Willison
View Original Article
Share this article:
Summary & Key Takeaways
- A researcher discovered a prompt injection attack against Claude Code Opus 5's Auto Mode.
- The attack tricks Claude into downloading, uncompressing, and executing local code.
- Crucially, Auto Mode sometimes blocked Claude's attempts to terminate the malicious process.
- The findings reinforce the necessity of running AI agents within robust sandboxes.
Our Commentary
This is terrifying. Anthropic made bold claims about Auto Mode, and now we see it not only fails but actively prevents the agent from cleaning up its own mess. The idea that the safety mechanism itself becomes part of the failure is a nightmare scenario. I've been saying it for ages: sandboxing is non-negotiable for agents. This proves it.
View Original Article
Share this article: