Breaking Claude Code Opus 5 Auto Mode

(embracethered.com)

114 points | by Recursing 4 hours ago

13 comments

  • andai 16 minutes ago
    >But it runs that decoder inside the attacker-controlled directory (unzipped archive)

    >There a malicious struct.py shadows Python’s standard implementation

    I ran into this myself, where some file I had given a random name turned out to shadow some Python standard library module, giving me the weirdest startup crash ever.

    That definitely doesn't seem to me like how that should be designed, magically silently importing everything you see and overriding basic functionality.

    • DanielHB 5 minutes ago
      Claude ran npm update (update all dependencies to the latest version compatible with the semver specified) in my repo without telling me when trying to fix some problems. Given that only updates the dependency lock-file I didn't notice and it caused several hours of debugging for me.

      It is quite sneaky how LLM output can sometimes bypass human verification like that. No one is going around checking every single line change in auto-generated files. Someone could easily sneak a malicious dependency in there through some online tutorial that the LLM searches for.

  • rcxdude 51 minutes ago
    I would not really call this a prompt injection attack, since it doesn't really hijack the agent to become malicious (something the article does discuss later on). It's more a trojan that's aimed at tricking Claude specifically.
  • comboy 1 hour ago
    Interesting attack, very nicely designed. Not sure if it's much related to the auto mode itself though.
    • kevsim 1 hour ago
      The point is that auto mode gives people a false sense of security that leads them to believe they don't need to run Claude in a proper sandbox. This same attack running in a sandbox (even in YOLO mode) would be comparatively harmless.
      • bombcar 0 minutes ago
        Auto mode is for people who just keep hitting "YES" on everything, it's a bit better than that.

        But it's real easy to give auto mode instructions (like "always ask before deploy") and then bypass that just normally.

    • yeputons 56 minutes ago
      I don’t think it’s related to the auto mode at all. It would work perfectly in the manual mode. It does not even need Claude: just give a human a similar archive and hope they run some simple Python from the directory at least once. And make sure there are lots of files do they don’t notice a weird .py around
      • rcxdude 46 minutes ago
        Probably the human would just run the binary.
  • colinmarc 47 minutes ago
    What's interesting to me about this is that it targets Claude's specific tics. Anthropic has created model that reliably reaches for the same tools (yes, and phrases; `python -c` is a load-bearing tool for it). Everyone gets the same model, so by learning the model's behavioral patterns you can target it better.
    • throwawayffffas 33 minutes ago
      Kimi and GLM reliably run python to do stuff as well.
      • andai 14 minutes ago
        If my memory serves me so do GPT and Deepseek. So I'm not sure if this attack is Claude specific at all.
  • too_pricey 40 minutes ago
    As discussed [here](https://lobste.rs/s/ktbweg/prompt_injection_claude_code_opus...), this is not prompt injection. The prompt was to summarize the website, and in the process of summarizing the website, Claude writes a decoding script that it runs in an insecure and exploitable way. At no point was the intent of the agent hijacked, this was just code. Which is potentially more interesting!
    • rcxdude 37 minutes ago
      It does mean that you could potentially hijack the agent afterwards, though, which could make the trojan into an even bigger threat.
      • lnenad 33 minutes ago
        But you're not actually hijacking the agent if you start a new process.
        • pests 19 minutes ago
          The agent wrote the code that triggered a vuln and allowed you to start the process
    • lnenad 33 minutes ago
      Yeah, I agree, this is a different vector. Still scary though and very related to AI.
  • hahn-kev 54 minutes ago
    As a non Python dev this seems like very surprising behavior for a system library to be modified by just having a file with a specific name in the same folder.
    • rcxdude 50 minutes ago
      You would get a similar thing in C and C++ with a system header in a library directory (maybe some compilers would warn on such a thing?). Most languages don't privilege their standard libraries in a way that would prevent this.
  • Phemist 1 hour ago
    This default-to-auto-mode and the misleading marketing is begging for a class action once damages accumulate. Especially considering the Auto Mode even can actively prevent the clean-up!
    • thewhitetulip 44 minutes ago
      Well, the jokes on us because laws don't apply to AI firms
  • nasretdinov 1 hour ago
    That's an interesting technique! I'd also like to point out that there's something odd with the page itself too, my phone got really hot while I was reading the page, and drained a significant amount of battery charge as well.
  • mcherm 48 minutes ago
    Why is the first step needed? What does the use of WGet (rather than curl) do to block this attack?
    • rcxdude 35 minutes ago
      They only mention it in passing, but I think it's mainly just the default tool call (which isn't wget, it's a built-in thing in the harness) just throwing off Claude's habits a bit (and not always just downloading the file).
  • throwawayffffas 39 minutes ago
    > In a few runs Claude tried to terminate the malware process once it noticed the compromise, but Auto Mode denied the cleanup command.

    See, that's why you should run with --dangerously-skip-permissions

    Jokes, aside running with dangerously-skip-permissions is really handy, and I have found that I cannot be trusted to vet commands and code, and guess that automode is only marginally better than a human, and the cost of false positives is too high for my workflow.

    So skipping permissions is where we are at, and disallowing network access seems to be the way to go.

  • julien_dev 2 hours ago
    I'm quite surprised that we are not seeing something like this more in the wild. Quite concerning
    • chmod775 22 minutes ago
      What vibe "don't even gotta read it" coder would even notice if it happened?
    • mkurz 1 hour ago
      Maybe it is used in the wild, but we just don't know.
    • rcxdude 59 minutes ago
      It's not that far off a typical trojan, just one tailored to Claude's habits. A lot of the same limits apply.
  • bewareofscams 1 hour ago
    > Boris Cherny from Anthropic recently posted that layered defenses could reduce indirect prompt injection on unseen attacks to approximately zero.

    > I got attack success rates up to 80% using a small sample size.

    Snake oil salesman misrepresents the data. Color me surprised! /s