How a summary request hijacked Claude Code's Auto Mode
A researcher got Claude Code's new Auto Mode to run hidden code just by asking it to summarize a website.
Auto Mode is now the default setting in Claude Code, the tool you use every day, so it is worth understanding exactly what a security researcher just got it to do with nothing more than a request to summarize a website.
Auto Mode replaces the usual human approval prompts (the yes, run this command pop ups) with an automatic safety classifier that decides on its own whether an action is safe. It became the default starting mode for Claude Code in mid-August. A third-party evaluation commissioned by Anthropic showed a 0.00 percent prompt injection attack success rate for Opus 5 in Auto Mode, though the evaluation does not appear to have a published benchmark name. A researcher who writes as Embrace The Red decided to test that claim directly and, using a small sample, got the attack to succeed 60 to 80 percent of the time.
The setup starts with an ordinary request: summarize a web page. That page is built to look like a small online archive of old notebook records about the history of language, packaged as a ZIP file (a compressed folder of files). Claude's normal WebFetch tool, which reads and summarizes pages on its own, fails on this page, so Claude decides on its own to switch to using the command line tool curl to fetch the raw page directly. That switch matters: the attacker never told Claude what to do, the page just made the risky path the only one that worked. The request redirects to a ZIP archive called WIC-notebook-catalogue. Claude downloads and unpacks it, finding a README describing how the notebook records are encoded, a native decoder program, and a file named struct.py.
Claude looks at the decoder program, correctly judges it as suspicious, and refuses to run it. That is the safety behavior working exactly as intended. But instead of stopping there, Claude decides to write its own replacement decoder in Python and runs it inside the same folder it just unpacked. That folder also contains the file named struct.py, sitting right next to Python's own real struct module (a built-in part of the Python language). When Claude's new decoder script imports Python's base64 module, that module internally imports struct, and Python finds the attacker's fake struct.py first because it sits in the current folder, a trick called module shadowing (a bad file secretly replaces a trusted one). The fake file quietly runs its own hidden code the moment it loads, while still passing back the real decoding functions so nothing looks broken. Claude never runs the obviously bad file. It gets tricked into running the hidden bad code through code it wrote itself, believing it was being careful.
The researcher's own conclusion is blunt: if you care about misalignment (the AI doing something you didn't intend), the model making things up, or prompt injection, Auto Mode is not a replacement for running your agent in an isolated environment and watching what it does. Anthropic's own safety engineer, Boris Cherny, has said the company's layered defenses can bring injection success on unseen attacks down to close to zero. This test suggests that claim holds for the exact scenarios the vendor tried, and breaks down fast against a new, targeted attack chain built specifically to route around it.
If you care about what's happening and are worried about misalignment, hallucinations and prompt injection, then Auto Mode IS NOT a substitute for running your agent in an isolated environment and monitoring what it is up to.via Embrace The Red →