A plugin's pinned commit can be swapped for attacker code in four AI coding agents
AIR Security reports that Claude Code, Codex, GitHub Copilot and Gemini CLI each check out a plugin's pinned commit without confirming the checkout landed there — “That one missing check is the whole bug” — so an attacker controlling a plugin repository can substitute code that background auto-updates then install without user interaction. Anthropic fixed it in Claude Code 2.1.179 and OpenAI in Codex 0.146.0; Microsoft has shipped no fix for GitHub Copilot and Google deprecated Gemini CLI rather than patch it. The research was found in May 2026, disclosed to the four vendors in June, and carries no CVE identifier.
A study of twelve agent harnesses finds attacker text being promoted into a higher-privileged message role
A preprint presents what its authors call “the first systematic analysis of context assembly designs in real-world AI agent harnesses,” naming two classes of flaw: MessageRole Context Privilege Escalation, in which “attacker-controlled content originating from a low-privileged context is incorporated into a higher-privileged message role,” and Cross-Scope Context Privilege Escalation, which “occurs when attacker-controlled content persists beyond the context in which it was introduced.” The paper reports testing “12 real-world agent harnesses, including Claude Code and Codex,” with consequences it lists as “full agent compromise, remote code execution, denial of service, and manipulated tool or skill invocations.”
A repository's own git config makes seven AI coding agents run attacker code before any prompt
Manifold Security reports eight findings across seven AI coding agents in which a repository's git configuration names a command that git then executes on the host, with the user's privileges, before any trust prompt, because agents run git commands at session start to gather context. The named vector is the core.fsmonitor setting; Claude Code, Goose, OpenAI Codex and Cursor shipped fixes while Qwen Code, Grok Build, Hermes Agent and a second Claude Code path were unpatched at publication. The write-up states two CVEs, CVE-2026-72718 for Goose and CVE-2026-71963 for Hermes, and says delivery requires the repository to arrive as files with its .git directory intact rather than through a clone.
Anthropic ships Fable 5.1 generally and keeps Mythos 5.1 behind trusted-access vetting
Anthropic says Mythos 5.1 “demonstrates the strongest cyber capabilities of any model we've released” and is available only through its trusted access programs, while Fable 5.1 is generally available. It says Claude Code users can expect “an average of around 60% fewer interventions per session from our cyber safeguards” relative to the previous safeguards on Fable 5, with dual-use tasks including penetration testing, exploit generation and binary-based vulnerability scanning still routed to Opus models.
Researcher reaches code execution in Claude Code's Auto Mode by shadowing a Python module
Johann Rehberger redirected Claude from its WebFetch tool to curl using an HTTP 415 response, served a ZIP archive containing a malicious struct.py, and obtained remote code execution when Claude's own decoder imported a module that in turn imported the shadowed one — reporting a 60 to 80 percent success rate across payload variants on small samples. Anthropic closed the report as “Informative,” saying Auto Mode is a convenience feature backed by a best-effort classifier rather than a security guarantee; Rehberger notes his chain was not among the 72 scenarios behind a previously cited near-zero prompt-injection figure.
Rapid7 finds a crypto-fraud crew used Claude Code to build and run a vishing pipeline against wallet users
Rapid7 Labs, analysing an exposed web directory and recovered session logs from a cryptocurrency fraud operation it named ASTERIX, found the operators used Anthropic's Claude Code to manage target lead lists and configure network infrastructure — cleaning a dataset of more than 103,000 Polish phone numbers and setting up scripts to validate numbers against Crypto.com and Kraken accounts — as part of a pipeline of phishing, vishing and fake wallet apps built to steal recovery phrases. The exposed server held roughly 885,000 phone numbers across 54 countries. When the operator asked Claude to help obfuscate a malicious build, Claude declined, and the operator switched to Moonshot's Kimi model with a jailbreak prompt.
Poisoned observability logs drive AI coding agents, with a sandbox escape patched before disclosure
Tenet Security reports that error and observability data from services such as Sentry, Cloudflare and Datadog can act as an indirect prompt-injection channel into AI coding agents, claiming a 90% success rate against Claude Code running Sonnet 4.6 in Cloudflare's recommended setup, and estimating more than 15,000 organisations exposed by extrapolating from 73 public artifacts across 48 organisations. Anthropic confirmed and fixed a Claude Desktop sandbox escape used in the chain before publication, with no CVE assigned; Sentry, Datadog and Cloudflare were notified between June 3 and July 13.
Cisco Talos analyses prompt logs recovered from threat actors' own machines
Talos examined a corpus of prompt logs left by Claude Code, CodeX, Cursor and Gemini on threat actor endpoints, grouping the use into AI as a malicious software engineer, AI for scaling criminal operations and AI for vulnerability research. It reports it “did not encounter any sophisticated encoding or techniques designed to trick the models” — claims of equipment ownership, capture-the-flag or bug-bounty framing, splitting risky actions across sessions and neutral verb choice were enough — and concludes “guardrails are not functioning as expected.”
npm worm in keyv and cacheable namespaces steals AI coding-tool credentials and persists via Claude Code and VS Code hooks
A self-propagating npm supply-chain compromise spread from the keyv and cacheable namespaces into over 400 packages, using a preinstall script to harvest cloud credentials, CI/CD secrets, private keys and cryptocurrency wallets, and republishing poisoned versions through npm OIDC trusted publishing. The payload specifically targets Claude, OpenAI, Codex, Cursor and Gemini credential stores and plants autostart hooks in .claude/settings.json and .vscode/tasks.json so that the payload runs when a developer or an AI coding agent opens the cloned repository, with no npm install required.
Black Hat USA 2026 vendor announcements centre on AI agent runtime protection, discovery and least-privilege enforcement
SecurityWeek's three-part roundup of Black Hat USA 2026 announcements documents a concentrated wave of defensive products aimed at securing AI agents, including Cyera Agent Guardian and Menlo Security MARS for prompt-injection and exfiltration protection, KnowBe4 Agent Risk Manager and Mimecast Agent Risk Center for agent discovery and behaviour monitoring, Varonis intent-based access control and Zero Networks least-agency enforcement for constraining agent permissions, and Acalvio Deception Guardrails for honeytokens targeting agentic environments. Legit Security's VibeGuard 2.0 and Sysdig Secure AI specifically target AI coding agents such as Claude Code, Cursor and GitHub Copilot.
Hunt.io reports suspected China-linked operators running Claude Code and DeepSeek as an intrusion toolchain against government targets in four countries
Hunt.io analysed an exposed open directory containing 2,431 files, including operator logs of LLM sessions, and described a split-model workflow in which Claude Code acted as the execution engine for agentic tool use, bash execution and session persistence while DeepSeek-v4-pro handled attack logic, script generation and decision-making. Hunt.io reported exploitation against an Afghan government application, a Thai government administrative system via SQL injection, and two Taiwanese critical-infrastructure organisations, with reconnaissance against US government entities and financial firms in Europe, Australia and Asia.
Agent skill metadata fields can suppress permission prompts and hide a skill from the user
HiddenLayer reports that a Claude Code skill's allowed-tools frontmatter field bypasses permission requests for tools including Bash, that setting user-invocable to false keeps a skill out of the menu while leaving it available for background use, and that project memory files can be written without a permission request. It also shows a denial-of-wallet path, with one URL-summary task costing $0.0274 on a small model at low effort and $0.1451 when the skill specifies a larger model at high effort. The write-up records no vendor acknowledgement or fix.