Breaking Claude Code Opus 5 Auto Mode
dev_tools
A security researcher has discovered a significant flaw in the auto mode of Claude Code, a coding agent from Anthropic. As Simon Willison's Weblog reports, researcher Johann Rehberger found a prompt injection attack that works eighty percent of the time by tricking the agent into executing malicious local files. In some instances, the safety mechanism actually blocked the agent's own attempts to terminate the malware once the compromise was detected. The findings suggest that the only secure way to run such agents is within a sandbox or virtual machine to restrict access to sensitive credentials.
Source: https://simonwillison.net/2026/Aug/27/breaking-claude-cod...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton