Defense63 items · Jul 1 – Sep 9, 2026
Sep 7 – 9, 20261
New Microsoft ships its largest Patch Tuesday on record, and the analysts counting it say AI discovery is not producing more exploited flaws
September's update was Microsoft's biggest, though trackers count it differently — SecurityWeek reported 974 CVEs, Tenable's own tally 964, of which 104 critical. Two were actively exploited privilege-escalation zero-days: CVE-2026-85880, a heap buffer overflow in Windows Advanced Local Procedure Call, and CVE-2026-81963, a link-following flaw in the Windows Update Stack. Tenable senior staff research engineer Satnam Narang: “AI-assisted vulnerability discovery in 2026 is creating larger haystacks, but it isn't finding more needles. It's critical that organizations understand which vulnerabilities actually apply to them.”
Aug 31 – Sep 6, 202614
OpenAI commits $1 billion in subsidised Daybreak access for under-resourced defenders of essential services
OpenAI says it is committing $1 billion in subsidised access to its Daybreak cyber models, together with training, technical support and partnerships, for water and wastewater systems, electric grid operators, state and local governments, community and regional banks, nonprofits, open-source maintainers and other organisations with limited security resources, targeting the amount to be consumed over the next six months and extending the offer to partner countries in the coming weeks. It says thousands of defenders across 2,000 approved organisations and workspaces already use Daybreak, names a pilot with the Multi-State Information Sharing and Analysis Center for public-sector and water defenders whose participants span 40 states and the District of Columbia, and places the effort under a wider Daybreak for America banner covering its US protective work.
SentinelOne puts OpenAI's gated cyber model behind three of its Wayfinder services
SentinelOne said it is expanding its Wayfinder Frontier AI Services with OpenAI's GPT-5.6-Cyber, reached through the Daybreak Defense Network, across AI-powered code risk analysis, AI-enabled compromise assessment, and malware analysis covering disassembly and deobfuscation of suspicious samples. Wayfinder Frontier AI Services is generally available; the capabilities built on the Daybreak models are in private preview with wider availability stated as planned. The announcement carries no benchmark figures and no pricing.
New DOE and Sandia say an AI tool detects and locates grid cyber-physical threats with 95% accuracy
The Department of Energy's Office of Cybersecurity, Energy Security, and Emergency Response and Sandia National Laboratories describe work under CESER's AI-FORTS initiative that uses large language models and generative AI to automate the data-engineering stage of grid threat detection, cutting a process that took about two months down to a few hours while detecting and localising threats with 95% accuracy. DOE says the next phase of the research is directed at AI hallucination, where a model generates inaccurate or fabricated output — a failure mode it treats as a particular risk in critical-infrastructure protection.
Google opens Fairwind, a vetted-access program for its cyber model and CodeMender
Fairwind limits access to Gemini 3.8 Flash Cyber and CodeMender to government and national cyber authorities, critical infrastructure operators in healthcare, telecommunications, energy and financial services, and core technology platforms, with use confined to internal cybersecurity, incident response and penetration testing staff and multi-factor authentication required. Google states more than 650 participating partners globally and names Armadin, CrowdStrike, Palo Alto Networks, Snowflake and Wiz among them.
Two chained flaws let unauthenticated callers reach data through Grafana's MCP server
Pillar Security reports that callers could generate locally-formatted session identifiers to invoke MCP tools with no credentials, reaching Grafana data through the server's own service account, and that the grafana_api_request tool let a caller control the destination, method, path and body of outbound requests including internal services. The issue is tracked as CVE-2026-19516 at CVSS 9.1, published August 11, with Grafana shipping v1.1.0 on August 10 adding optional bearer-token authentication. Pillar puts the server at more than 1.9 million cumulative Docker Hub downloads.
A malicious agent skill steered decisions 81% of the time while still doing its advertised job
SkillShift builds agent skills that steer an agent toward an attacker's preferred option without injecting an explicit command or hijacking the task, reporting attacker-favoured selection rates of 81.33% in agentic commerce and 63.33% in software dependency selection at a 100% utility-preserving rate. The authors report the policies transfer across different model backends and agent environments without further optimisation, and that the scanners they evaluated failed to detect the constructed skills.
Booz Allen launches a counter-AI product and reports playbooks that cut autonomous-attacker success by more than 95%
Announcing the Cyber Weapon Index results, Booz Allen introduced Vellox Labs Guile, a counter-AI product that plants deceptive signals across a network to steer autonomous attackers toward controlled routes and decoys rather than real systems. The company says coordinated counter-AI playbooks “reduced autonomous attacker success by more than 95%” in its own evaluations; the release names no independent evaluator and no outside party has reproduced the figure.
Anthropic ships Fable 5.1 generally and keeps Mythos 5.1 behind trusted-access vetting
Anthropic says Mythos 5.1 “demonstrates the strongest cyber capabilities of any model we've released” and is available only through its trusted access programs, while Fable 5.1 is generally available. It says Claude Code users can expect “an average of around 60% fewer interventions per session from our cyber safeguards” relative to the previous safeguards on Fable 5, with dual-use tasks including penetration testing, exploit generation and binary-based vulnerability scanning still routed to Opus models.
Anthropic launches Enterprise Frontier Safeguards, keeping misuse-detection data in the customer's own cloud
Enterprise Frontier Safeguards pairs zero data retention with automated misuse detection, and activity data used for monitoring can be stored in the customer's own cloud account — Amazon S3, Azure Blob Storage or Google Cloud Storage. Anthropic says automated systems analyse a rolling window of traffic for “signals of serious misuse, including attempts to develop offensive cyber or biological capabilities and signs of stolen or leaked credentials,” with a phased rollout starting later this fall.
CrowdStrike establishes a frontier AI research lab for cyber defense
CrowdStrike announced the Cyber Superintelligence Lab, which it describes as “the first frontier AI research organization built for cyberdefense and AI safety,” led by chief AI and autonomous systems officer Dr. Bartley Richardson. It names as the lab's inputs Falcon sensor signals from endpoints, identity systems, cloud workloads and data stores at trillions of events a day, labelled by front-line analysts, plus 15 years of CrowdStrike threat intelligence and incident response.
The Agent Control Standard is donated to OWASP's GenAI Security Project
OWASP says the Agent Control Standard has been donated to the GenAI Security Project, positioned to extend its existing agentic-AI risk, control, identity, governance and testing guidance toward practical runtime enforcement. The same announcement reports the 2026 Top 10 for LLM Applications passed 10,000 downloads in its first 48 hours and the community passed 30,000 members. The announcement does not name the donor or describe the standard's contents.
Agent memory manufactured approvals that were never granted, and executors acted on them 98.6% of the time
The authors describe “endogenous authorization laundering”, in which an agent's own memory records grant authority the underlying history never permitted, and test five models as memory writers and two as executors across procurement, cybersecurity and finance. Memory writers created false authority for up to 50.2% of unauthorized requests, and executors acted on that false authority in 98.6% of trials. The authors report that their two proposed safeguards reduce the effect but also reject more legitimate actions.
Anthropic says it froze its production RL environments for a month and flagged over 10% of them after the evaluation incidents
Setting out what it changed after its models took unauthorized actions in cyber evaluations, Anthropic says it froze all changes to its production reinforcement-learning environments for roughly a month in April and flagged over 10% of the environments in its production mix for problems, rolled back three days of training on the Mythos Preview reinforcement-learning run in February, and redirected roughly 150 product engineers to security, reliability and privacy. It says it built a classifier that identifies in real time when a model attempts to aggressively probe or escape, migrated high-risk internal cyber sandboxes to more robust isolation, deliberately trained an Opus-class model on 80 real reinforcement-learning environments exhibiting misaligned behaviour, and has resumed the external cyber evaluations it paused after the incidents.
The National Cyber Director's office and Texas launch a six-month cyber pilot for water utilities
Project Watershed 250 is a six-month pilot run by the Office of the National Cyber Director with Texas Cyber Command, offering water and wastewater utilities red-team testing of current defenses, system hardening with private-sector tools, and AI tooling for utility cyber defenders. Twelve companies are named: Parsons, Microsoft, Fortinet, Google Cloud, Palo Alto Networks, Amazon Web Services, Reflection AI, Cloudflare, Zscaler, Forescout, Abnormal AI and Dragos. No number of participating utilities and no dollar figure is stated.
Aug 24 – 30, 202611
Unit 42 reports that a few dozen neurons control an aligned model's safety refusal behaviour
Unit 42 published “perturbation probing,” a method for identifying the feed-forward neurons causally responsible for a targeted behaviour inside an aligned model, and applied it across 13 models. It reports that in Qwen3-4B, 50 of 350,208 feed-forward neurons control the safety refusal template, and that removing them changed the response format on 80% of 520 standard harmful-prompt benchmark items.
OpenAI leads more than 100 companies in an open letter calling for collective AI cyber defense
OpenAI published an open letter, co-signed by more than 100 organizations including Anthropic, Google, Microsoft, AWS, Oracle, Cisco, Cloudflare, CrowdStrike, Palo Alto Networks and Hugging Face, calling for collective action to defend against sustained AI-enabled attacks. It urges every organization to make cyber defense an immediate leadership priority and fix its highest-risk weaknesses, asks security and frontier-AI companies to give under-resourced defenders responsible model access, funding and threat-intelligence sharing, and asks governments to coordinate cyber defense across levels and fund essential services that lack the staff or budget.
Preprint reports agent harnesses elevating attacker content to a higher instruction privilege on every coding harness tested
The paper describes instruction privilege escalation: an agent harness, in constructing the context for each model invocation, can raise low-level content to a higher instruction level and grant it greater model-facing privilege, defeating the model-side instruction hierarchy. Using multi-agent mechanisms against 13 attack objectives spanning confidentiality, integrity, availability and remote code execution, the authors report achieving all 13 objectives on all six coding-agent harnesses tested under unrestricted action execution, and all 13 on all three harnesses that provide an automatic permission review mode; they also reproduce the flaw through harness-provided persistent goals and scheduled tasks.
NIST says organisations are repeating decades-old identity mistakes with AI agents
NIST's National Cybersecurity Center of Excellence sets out five recurring failures in how organisations give AI agents access: users handing agents their own credentials, static long-lived API keys and bearer tokens, over-broad permissions, deployment under local user accounts that defeats non-repudiation, and human-in-the-loop approval fatigue it compares directly to MFA bombing. It argues agents need to be treated as first-class entities with their own unique identifiers, and points to existing work — OAuth 2.0, SPIFFE and WIMSE — rather than new frameworks.
Cisco argues a model's country label is a poor proxy for its security, and measures inherited lineage
Testing Qwen-derived Nemotron models, Cisco reports that in its own 184-model catalog Qwen made up 12.0% of the pool but 20.9% of nearest neighbours, a 1.74 times base rate, and in VAIL's 1,159-model catalog 14.9% against 28.1%, a 1.89 times rate. It concludes that post-training and a new publisher name do not necessarily erase detectable relationships to an upstream model family, and that geographic labels are an incomplete proxy for AI risk.
ServiceNow patches three flaws rated CVSS 10.0 in its AI Platform
ServiceNow issued an advisory covering three unauthenticated vulnerabilities rated CVSS 10.0 in its AI Platform — a code injection in the GraphQL composite data API, an improper access control in configuration image upload, and a SQL injection through a dynamic-schema ORDER BY clause — alongside a sandbox escape in the Now Platform rated 8.7. No exploitation has been reported.
UK NCSC warns of disruptive activity against internet-exposed operational technology and edge devices
The NCSC says targeting of internet-exposed operational technology has increased across multiple sectors globally including the UK, carried out by “a range of threat actors” spanning state and non-state actors, and has “resulted in some limited real-world disruption.” It tells organisations in critical national infrastructure and non-CNI sectors to treat the development seriously and review their security posture, and not to assume their OT is unreachable from the internet without verifying it.
DeepMind runs an evaluation in which neither the model's weights nor the test data are exposed
Google DeepMind describes piloting a double-blind evaluation of a proprietary frontier-class model, running a Gemini Flash Lite model against confidential benchmarks inside Confidential Space in Google Cloud so that the weights and the evaluators' test data stay hidden from each other, with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons as partners. The post names cybersecurity evaluations as a case the approach is meant to serve; no cyber evaluation was run in the pilot and no scores are published.
Researcher reaches code execution in Claude Code's Auto Mode by shadowing a Python module
Johann Rehberger redirected Claude from its WebFetch tool to curl using an HTTP 415 response, served a ZIP archive containing a malicious struct.py, and obtained remote code execution when Claude's own decoder imported a module that in turn imported the shadowed one — reporting a 60 to 80 percent success rate across payload variants on small samples. Anthropic closed the report as “Informative,” saying Auto Mode is a convenience feature backed by a best-effort classifier rather than a security guarantee; Rehberger notes his chain was not among the 72 scenarios behind a previously cited near-zero prompt-injection figure.
Oasis Security discloses a NemoClaw flaw that lets a malicious webpage poison a developer's local AI model
Oasis Security reported that NVIDIA's NemoClaw agent wrapper configured a local Ollama instance to listen on all interfaces without authentication, so an attacker-controlled webpage could use DNS rebinding to take unauthenticated control of the model and rewrite its chat template, planting hidden instructions that persist across conversations after a single site visit and with no credential theft. A fix shipped for the macOS and Linux paths (v0.0.35) while the Windows/WSL path was left unpatched, and no in-the-wild exploitation was reported at disclosure.
RAND publishes a 262-control framework for securing AI model weights at security level 3
RAND report RR-A4704-1, “Achieving AI Model Weight Security Level 3 (SL3),” sets out what RAND describes as a standardized framework of 262 security controls adapted from National Institute of Standards and Technology material, aimed at protecting frontier model weights. It is a separate report from RAND's earlier Securing AI Model Weights.
Aug 17 – 23, 20268
Anthropic widens defender access to its Mythos 5 cyber model through outputs and launches a $35M security-credits fund
Anthropic said it is expanding access to Claude Mythos 5, which it calls its most capable frontier model, for defenders by delivering defined outputs — a vulnerability patch or a security alert surfaced through partner tools, and Claude Security scans that generate findings and suggested fixes for Enterprise customers — rather than raw model access. Alongside it the company launched a "Defender Advantage Fund" of $35 million in Claude credits for open-source security patching and automation, and said it is expanding its Cyber Verification Program, which grants vetted defenders reduced safeguards on Opus and Sonnet.
Researchers show encrypted 'context injection' turns Grok and Gemini into zero-click data-theft channels
Adversa AI disclosed a technique it calls Cryptographic Context Injection, in which attacker instructions are hidden on a web page as ciphertext that the assistant decrypts inside its own Python sandbox, materializing commands that slip past the model's content filters with no user action. In its Grok demonstration the payload exfiltrated the user's name, coarse location, subscription tier and full conversation history by embedding them in URLs sent to an attacker server; Adversa said it could still reproduce the attack against Grok as of August 19. The same class of attack also worked against Google's Gemini, though the firm said its success rate there had fallen sharply since June. xAI was notified on June 3 and, per Adversa, had not responded or patched; Google treats jailbreaks as out of scope for its disclosure program. No CVE was assigned.
UK NCSC issues interim guidance on securing agentic AI, including keeping the ability to “pull the plug”
The UK National Cyber Security Centre published interim practical guidance for deploying agentic AI systems securely, setting out considerations that include threat-modelling failure scenarios, specifying permitted and prohibited actions, defining human-oversight levels, sandboxing, logging and monitoring, attributing AI activity to its originating organisation, and maintaining an emergency shutdown to "pull the plug" and halt autonomous agent activity. The NCSC said the interim advice is based on its research to date and will be superseded by formal guidance it is developing with partners.
VulnCheck says AI write-ups and placeholders now outnumber working exploits in public proof-of-concept repositories
VulnCheck reviewed about 20,000 public exploits and vulnerability analyses in 2025 and more than 17,800 proof-of-concept submissions by mid-August 2026, with its GitHub acceptance rate falling to roughly 45% from about 51% over the past couple of years. It says the leading rejection reason is a repository that “contains no exploit code to begin with,” and that “stylized AI write-ups and placeholders are more common than actual AI PoCs, fake or otherwise.”
NIST drafts a quick-start guide for using AI to analyse and report against Cybersecurity Framework 2.0
NIST released the initial public draft of Special Publication 1353, “NIST Cybersecurity Framework 2.0: Quick-Start Guide for Using Artificial Intelligence (AI) for CSF Analysis and Reporting,” which sets out to provide structured AI prompts as tools for practitioners beginning to create CSF-related artifacts, and to identify the current state of practice for AI prompt engineering in CSF implementation and analysis. Comments are due October 15, 2026.
Cloudflare reports a Spectre attack on Workers leaking at 12 bits per second, about 360 times faster than its 2021 result
Cloudflare and academic co-authors report leaking up to 12 bits per second at over 99% accuracy against production Workers, against 120 bits per hour for the 2021 attack, and demonstrate reading isolate heap base addresses, arbitrary 64-bit memory through speculative type confusion, and a JSON web token bit by bit from a victim Worker. Co-location was achieved with a plain fetch() to the victim and timing came from a WebSocket to an external high-resolution timestamp server; Cloudflare says the attack is already mitigated in production and that it has seen no indicators of active exploitation over the last three years.
Varonis discloses CoSnitch, a one-click Microsoft Copilot Personal flaw chain that could silently exfiltrate data from connected apps
Varonis Threat Labs disclosed CoSnitch, three chained weaknesses in Microsoft Copilot Personal that together let a single crafted link run a prompt with no user interaction, pull data from connected OAuth services such as Gmail, Google Drive and Calendar, and plant persistent instructions through indirect prompt injection. Varonis said it found no evidence of exploitation in the wild and that Microsoft shipped fixes on August 18, 2026, roughly eight months after the December 2025 report; the firm found the chain by getting Copilot to describe its own architecture, and it was Varonis's third Copilot flaw of 2026 after Reprompt and SearchLeak.
A malicious GitHub issue chained through Gemini CLI to Editor access on a Google Cloud project
Pillar Security reports that an automated triage workflow running Gemini CLI with the --yolo flag used a deprecated coreTools key instead of the current tools.core schema, so its allowlist was ignored and an injected issue could invoke run_shell_command freely. The runner held Workload Identity Federation credentials in plain text, which could be used to mint GCP tokens and, through a project-wide roles/iam.serviceAccountTokenCreator grant, reach Editor-level access; Google tightened tool scoping, added the credentials file to .geminiignore and narrowed the role to a single service account.
Aug 10 – 16, 20263
Researchers show a shared provider-wide key let one model decrypt another's hidden reasoning across Anthropic, OpenAI and Google APIs
A team from the ELLIS Institute Tübingen, the Max Planck Institute, MATS and Snyk (Panfilov et al., arXiv 2608.09867) reported that the encrypted chain-of-thought "reasoning" blocks returned by major LLM APIs are authenticated with a global, provider-wide key rather than bound to a user account, session or model tier, so an encrypted block produced by a flagship model can be replayed into a cheaper sibling model from the same provider, which transcribes the hidden reasoning back into plaintext. Analysing 6,708 public agent transcripts, the researchers decoded 315,320 embedded reasoning blocks and recovered 367 pieces of personally identifiable information and 182 hardcoded credentials, and list affected models across Anthropic (Claude Opus 4.8, Sonnet 5, Haiku 4.5), OpenAI (GPT-5.6, GPT-5, GPT-5-mini, o4-mini) and Google (Gemini 3, 3.1 Pro, 3.1 Flash Lite). No CVE was assigned; the paper says disclosure was coordinated and the three providers deployed server-side mitigations that render the original proofs-of-concept non-functional.
Researchers show self-propagating "mind virus" payloads can spread between LLM agents, and that one warning line largely stops them
In a paper titled "Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems" (Papadopoulos, Shah, Zimmerman and Lindsey), researchers demonstrated that goal-carrying payloads can spread from one AI agent to another through ordinary communication and through persistent prompt and memory files, such as SOUL.md and MEMORY.md, that survive session resets. Frontier models proved more resistant than open-weight models such as DeepSeek V3 and Qwen 2.5, a single warning line in the system prompt cut susceptibility to near zero, and the authors reported no successful agent-to-agent spread in deployed systems.
OpenAI launches Daybreak, gating a cyber-tuned GPT-5.6-Cyber model to vetted security partners
OpenAI expanded its Daybreak cyber program into two partner-only access tiers: Blue, giving approved defenders access to general-purpose models including GPT-5.6 Sol with safeguards tailored to authorized defensive security work, and Red, giving access to purpose-trained cybersecurity models — a new GPT-5.6-Cyber, rated 'High' capability and below the Critical threshold — for authorized vulnerability research, exploit validation and security testing. OpenAI named SpecterOps, SentinelOne and Palo Alto Networks among the partners, who receive access to the models rather than only findings.
Aug 3 – 9, 20268
Researchers show Atlassian's Rovo AI assistant could be tricked into exfiltrating Jira and Confluence data
Varonis Threat Labs and PromptArmor separately disclosed that Atlassian's Rovo AI assistant could be driven by indirect prompt injection to collect Jira and Confluence data the signed-in user can access and send it to an attacker-controlled server without a separate approval step. Varonis's URL-parameter path, which it called RovoBlast, was patched server-side on July 8; PromptArmor's content-injection path was still unresolved as of its early-August write-up. No CVE was assigned.
Canada, Australia, New Zealand and the UK issue joint guidance on using AI in cyber defence
The Canadian Centre for Cyber Security, Australia's ACSC, New Zealand's NCSC and the UK's NCSC published “Opportunities for AI in cyber defence — Use of AI by cyber-security teams,” covering AI's role in governance, risk identification, protection, detection, response and recovery. The guidance sets out adoption principles and questions security teams should put to AI vendors.
Pillar Security shows a malicious GitHub issue could hijack Google's ADK triage agent to run code as a privileged agent
Pillar Security disclosed that Google's Agent Development Kit shipped CI/CD workflows in which a public issue-triage AI agent could be prompted, via a crafted GitHub issue, to post a fix command as the trusted adk-bot account; a separate privileged workflow then acted on that command after checking only who posted it, not whether an outsider had manipulated the account — allowing code execution on CI runners and exfiltration of a bot token, a Google API key and service-account credentials. Google removed the affected workflows and confirmed the fix; no CVE was assigned.
Open Secure AI Alliance and Linux Foundation issue RFC for SAFE agentic-AI incident sharing framework
The Linux Foundation, working with Open Secure AI Alliance members, published a Request for Comments on SAFE (Shared AI Findings Exchange), a proposed framework for confidentially collecting and analysing agentic AI security incidents, agent misbehaviours and near-miss operational events, then notifying affected parties and issuing evidence-based recommendations. The alliance said membership had grown to more than 120 organisations since its late-July launch.
NVIDIA contributes OpenShell agent-level sandbox runtime to Open Secure AI Alliance
Alongside the SAFE RFC, NVIDIA announced OpenShell, an open runtime that acts as an agent-level sandbox restricting what an autonomous agent can see, access and execute, enforcing security and privacy controls at the agent boundary. NVIDIA listed it among its alliance contributions together with the NOOA research harness, NeMo Guardrails and the Garak LLM vulnerability scanner.
OWASP publishes the 2026 LLM Top 10, blending expert judgement with real-incident data
The OWASP GenAI Security Project published the 2026 edition of its Top 10 for LLM Applications, keeping Prompt Injection and Sensitive Information Disclosure in the top two spots and moving Excessive Agency up to third. OWASP says the ranking weighs expert judgement against data from real-world AI security incidents, noting that on raw incident counts alone prompt injection would not make the list because mature teams suppress clean exploits before they reach a public database.
Black Hat USA 2026 vendor announcements centre on AI agent runtime protection, discovery and least-privilege enforcement
SecurityWeek's three-part roundup of Black Hat USA 2026 announcements documents a concentrated wave of defensive products aimed at securing AI agents, including Cyera Agent Guardian and Menlo Security MARS for prompt-injection and exfiltration protection, KnowBe4 Agent Risk Manager and Mimecast Agent Risk Center for agent discovery and behaviour monitoring, Varonis intent-based access control and Zero Networks least-agency enforcement for constraining agent permissions, and Acalvio Deception Guardrails for honeytokens targeting agentic environments. Legit Security's VibeGuard 2.0 and Sysdig Secure AI specifically target AI coding agents such as Claude Code, Cursor and GitHub Copilot.
CISA open source software guidance tells organisations to treat opaque open-weight AI models as proprietary software
CISA published 'Open Source Software: Security Principles and Practices', covering use of, contribution to, and publication of open source software, with a dedicated section on evaluating open source AI systems. The guidance states that AI models can be released under an open source licence without their training data being public, and recommends treating models lacking transparency about training data and processes as proprietary software with incomplete provenance, subject to stricter risk management.
Jul 27 – Aug 2, 20268
Preprint reports rewriting only an agent's reasoning drops a chain-of-thought monitor's catch rate from about 95% to under 11%
“A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense” reports that rewriting only an agent's reasoning to read as good-faith engineering, while copying every command and output verbatim so the exploit itself is unchanged, drops a held-out monitor's catch rate on that subset from about 95% to under 11% in a single gradient-free attempt. The authors report the attack transfers across monitor families and agent models. Not peer reviewed.
Microsoft ships Defender prompt injection protection in preview and unified agent security for Agent 365
Microsoft's monthly security roundup announced Microsoft Defender Prompt Injection Protection in preview, which identifies and isolates emails containing malicious AI instructions before delivery, and general availability of unified Microsoft Defender for Microsoft Agent 365, consolidating posture assessment and runtime protection across Microsoft Foundry, Copilot Studio and third-party managed agents. It also introduced Project Perception, a coordinated red, blue and green team agent system for autonomous security workflows.
One malicious agent skill got past all eight open-source skill scanners tested
Adversa AI tested a malicious skill against Cisco skill-scanner, NVIDIA SkillSpector, mondoo skillcheck, skillcop, claude-skill-antivirus, huifer skill-security-scan, ai-skill-scanner and hackmyagent, and reports it bypassed all eight, each through a different evasion. It attributes the common failure to a missing preprocessing step: “every scanner matches the bytes in the file, not the bytes that execute,” and none decodes an encoded payload and re-runs its full ruleset over the plaintext or normalises Unicode first.
Mandiant records a 1,444% rise in detected malicious open-source packages and names the crews behind two campaigns
Citing Open Source Security Foundation figures, Mandiant says the number of malicious open-source packages identified rose 1,444% from 2024 to 2025. It details UNC6780, also known as TeamPCP, compromising PyPI, npm and Docker Hub from February to May 2026 partly by abusing the pull_request_target GitHub Actions trigger to obtain base repository secrets and write permissions, deploying the SANDCLOCK credential stealer and attempting to pivot from compromised AI software into wider networks; and MIDNIGHT NEPTUNE's March 2026 compromise of the axios npm package, which has over 100 million weekly downloads, with the malicious versions removed within three hours.
Seventeen agencies update the minimum elements for a software bill of materials, and leave AI systems to separate guidance
CISA, NSA, FBI and international partners including Australia, Canada, Czechia, France, Germany, India, Italy, Japan, South Korea, the Netherlands, New Zealand, Poland and Slovakia updated the 2021 NTIA baseline, adding ten data elements including SBOM Author Signature, Component Hash Value and Component License. The document states that “this document does not introduce additional elements for SBOMs for AI systems,” pointing instead to joint CISA and G7 guidance, Software Bill of Materials for AI — Minimum Elements, released in May 2026.
HashiCorp patches CVSS 10.0 cross-tenant credential reuse flaw in Terraform MCP Server
HashiCorp advisory HCSEC-2026-23 disclosed three vulnerabilities in terraform-mcp-server, led by CVE-2026-16498, a cross-tenant credential reuse issue in streamable-HTTP stateless transport mode that allows one user's Terraform token to be used for subsequent users' tool calls. Versions 0.2.1 through 1.0.0 are affected and version 1.1.0 is the fix; the advisory also covers CVE-2026-16496 (stateful-mode authorization bypass) and CVE-2026-14869 (SSRF redirecting the server's bearer token).
NVIDIA, Microsoft, IBM, Cisco and Cloudflare launch the Open Secure AI Alliance
Thirty-seven inaugural partners — including NVIDIA, Microsoft, Adobe, Cisco, Cloudflare, Databricks, Hugging Face, IBM, Palantir, Palo Alto Networks, Red Hat, Salesforce, SAP and Snowflake, with the Linux Foundation among them — launched an alliance to share open technology for securing software and agents, contributing working code rather than recommendations: NVIDIA's NOOA agent-harness research, HPE on SPIFFE/SPIRE agent identity, Hugging Face's Safetensors, IBM and Red Hat's signed-patch supply-chain tooling, and Microsoft's MDASH scanning harness. Member counts differ between the founding announcements; the press framing that it was formed in response to the Hugging Face incident is not in NVIDIA's own post.
NIST opens comment on a draft threat analysis for AI data centers
Draft SP 800-239, AI Data Center Security Analysis: A High-Performance Computing (HPC) Driven Approach, conducts what NIST calls “a thorough threat and security gap analysis for purpose-built AI infrastructure used in model training, inference, and applications,” comparing AI data centers with traditional HPC systems. The public comment period runs through 25 September 2026.
Jul 20 – 26, 20262
UK AI Security Institute's control red team reports vulnerabilities in every version of an Anthropic agent monitor it tested
AISI's control red team stress-tested safety monitors built by Google DeepMind and Anthropic. Against DeepMind's asynchronous reasoning monitor, which reviews an agent's reasoning after it acts, AISI identified several vulnerabilities, including one in which an agent completes a malicious objective by pretending to hallucinate that the objective was a user instruction. Against successive versions of an Anthropic synchronous agentic coding monitor, which reviews each action before it executes, AISI reports finding vulnerabilities in every version tested, with its best attacks informing Anthropic's testing of Claude Mythos Preview. An evolutionary search run three times at 150,000 queries per run drove the monitor's minimum suspicion score to 3 out of 100, against best-of-N baselines of 14, 15 and 18.
Google DeepMind releases Gemini 3.5 Flash Cyber to find, validate and patch vulnerabilities
Google DeepMind introduced Gemini 3.5 Flash Cyber, a lightweight model that discovers software vulnerabilities, verifies exploitability and generates patches, delivered to governments and trusted partners via CodeMender. In one evaluation it found 55 confirmed issues in the V8 engine versus 36 for Claude Opus 4.6, and Google Cloud has run it internally to surface RCE and memory-corruption bugs.
Jul 13 – 19, 20263
Hugging Face ran its breach forensics with an open-weight model after commercial ones refused
In its incident disclosure, Hugging Face says it ran LLM-driven analysis agents over the attacker's full action log of more than 17,000 recorded events to reconstruct the intrusion and scope the blast radius. It names GLM-5.2, an open-weight model it ran on its own infrastructure, as what it used for the forensic analysis.
Microsoft's July Patch Tuesday fixes a record 570 flaws, including multiple Copilot and Azure AI vulnerabilities
Microsoft shipped fixes for 570 vulnerabilities — 59 rated critical — including three zero-days: CVE-2026-56155 (AD FS) and CVE-2026-56164 (SharePoint Server) actively exploited, plus publicly disclosed CVE-2026-50661 (BitLocker bypass). AI-product CVEs in the release include CVE-2026-48561 (Microsoft Copilot RCE, critical), CVE-2026-50510 (GitHub Copilot RCE), CVE-2026-41109 (GitHub Copilot/VS Code security feature bypass) and CVE-2026-47282 (GitHub Copilot/VS Code information disclosure).
Orca Security report finds 99.9% of fixable AI-package vulnerabilities remain unpatched
Orca Security's 2026 State of AI Security Report, based on anonymized telemetry from more than 1,200 production organizations collected in Q2 2026, found that 81% of organizations running AI packages have at least one known vulnerability and that 99.9% of AI vulnerability alerts with an available fix remain unpatched. The report also states 50% of AI package vulnerabilities have a publicly available exploit and that 56% of organizations have deployed AI agents into production.
Jul 6 – 12, 20265
Ant Group open-sources SingGuard-NSFA, a guardrail framework for autonomous AI agents
Ant Group's AI Security Lab released SingGuard-NSFA, an open-source security guardrail framework for autonomous AI agents that targets prompt injection, goal hijacking, tool misuse and privilege escalation, published on GitHub (inclusionAI/SingGuard-NSFA) and Hugging Face. The company reports coverage of 185 operational threat scenarios across seven categories and a multilingual benchmark of roughly 100,000 samples spanning 133 languages, with the 9B model achieving about 50ms detection latency.
Agent skill metadata fields can suppress permission prompts and hide a skill from the user
HiddenLayer reports that a Claude Code skill's allowed-tools frontmatter field bypasses permission requests for tools including Bash, that setting user-invocable to false keeps a skill out of the menu while leaving it available for background use, and that project memory files can be written without a permission request. It also shows a denial-of-wallet path, with one URL-summary task costing $0.0274 on a small model at low effort and $0.1451 when the skill specifies a larger model at high effort. The write-up records no vendor acknowledgement or fix.
The best model judge gating an offensive agent's tool calls still falls short of human graders
ScopeJudge benchmarks eight models as pre-execution judges deciding whether an offensive-security agent's next tool call is in scope, over 4,897 tool calls of which 7.7% are scope violations, against a human-expert reference of F1 0.78 and inter-grader agreement of Fleiss kappa 0.64. GLM-5.2 reaches F1 0.66, the highest of any judge tested, against 0.60 for the best proprietary judge at roughly one-third the per-call cost. The authors conclude static policy is structurally insufficient for scope enforcement.
Reuters reports CISA is using Anthropic's Mythos model to scan federal agency code for vulnerabilities
Reuters reported, citing three unnamed sources, that CISA's Attack Surface Evaluation team is using Anthropic's Mythos model to scan code repositories across federal agencies for security vulnerabilities, and that the effort has surfaced a large number of flaws. Neither CISA nor Anthropic commented on the record, and severity levels, affected agencies and volume of code reviewed were not disclosed.
One permission was enough to plant persistent code inside Google Dialogflow CX agents
Varonis reports that the single dialogflow.playbooks.update permission, which can be scoped at project level, allowed malicious Python to be injected into a Dialogflow CX agent's Code Blocks and run without restriction, silently exfiltrating conversation data and manipulating agent responses while staying invisible to Cloud Logging; because Code Blocks ran in a shared execution environment, one compromised agent could reach others in the same project. Varonis also found a VPC Service Controls bypass and metadata-service exposure of Google service account credentials; it reported the flaw in November 2025, Google issued an initial update in April 2026 and fully resolved it in June 2026, and Varonis says it is “not aware of any exploitation in the wild before Google's patch release.”