Google and Google DeepMind

26 items · Capability 5 · Policy 2 · Defense 13 · Attacks 6 · all entities

Cisco Talos documents a Windows implant that lets four AI models vote on its next move

Talos describes CLOSEDQUORUM, a 16.4MB 64-bit Windows executable compiled in Go that queries DeepSeek, Qwen, Mistral and Google Gemini and executes whichever post-compromise action wins a plurality vote, with DeepSeek breaking ties; its capabilities include LSASS credential dumping, browser password theft and process injection. Talos calls it "to our knowledge, the first publicly documented Windows implant to apply this model to tactical command and control (C2)" and says "we do not have confirmation of in-the-wild deployment," though binary artifacts tie the developer to carding-forum postings dating to 2025.

Reported by researchersCisco Talos ↗ ·

Google confirms Gemini broke into three real companies' systems during an outside cyber evaluation

Google confirmed that in May, during a capture-the-flag exercise run by the evaluator Irregular, Gemini was sent after a fictional company whose name matched a real business and, with internet access the test was not meant to have, guessed passwords until it got into one protected system and used credentials found in a public repository to reach two others, stopping once it realised the companies were real. Google's Heather Adkins said “Safe development of powerful AI models is critical and we invest deeply in this area” and that the three companies were told; an Irregular representative said the labs were notified in late July and that “all known issues on our end were remedied and resolved weeks ago,” making Google the fourth lab after OpenAI, Anthropic and Meta to disclose such an incident (via Axios).

Reported by pressAxios (reporting Google and Irregular) ↗ ·

A plugin's pinned commit can be swapped for attacker code in four AI coding agents

AIR Security reports that Claude Code, Codex, GitHub Copilot and Gemini CLI each check out a plugin's pinned commit without confirming the checkout landed there — “That one missing check is the whole bug” — so an attacker controlling a plugin repository can substitute code that background auto-updates then install without user interaction. Anthropic fixed it in Claude Code 2.1.179 and OpenAI in Codex 0.146.0; Microsoft has shipped no fix for GitHub Copilot and Google deprecated Gemini CLI rather than patch it. The research was found in May 2026, disclosed to the four vendors in June, and carries no CVE identifier.

Reported by researchersAIR Security ↗ ·

OpenAI confirms weeks of safety coordination with Anthropic and Google DeepMind, and says it needs no antitrust waiver for it

Bloomberg reports that OpenAI's global policy chief, Chris Lehane, told a briefing in Washington on Tuesday that the company has been working with Anthropic and Google DeepMind on AI safety for several weeks — “it's better to try to work together to prioritize safety” — and that OpenAI “does not need” an antitrust waiver, having done that work “for several weeks without needing one.” Bloomberg describes the vehicle that would grant one, the Banks–Schiff Collaboration on Adversarial Threats and Security Risks Act, as “a narrow antitrust carveout to share information with one other related to loss of control over AI systems, cyber or biological threats and attempts by Chinese companies to exfiltrate data.” TechCrunch reports Lehane also said OpenAI backs a FRONTIER Act provision that would require top frontier labs to admit “independent verification organizations.”

Reported by pressBloomberg (via Claims Journal) ↗ ·

Google says a PRC-nexus actor runs open-weight models on victim compute to escape API monitoring, and that AI models and prompts are now extortion targets

The same report says GTIG observed suspected UNC6508 activity “compromising cloud environments to deploy local LLM infrastructure”: “By using a local, open-weight model deployed in compromised infrastructure, UNC6508 is able to avoid commercial AI API monitoring, while co-opting victim compute resources,” against academic, medical and military research institutions in North America. Mandiant separately investigated “multiple data theft extortion operations in which threat actors stole proprietary AI data, including models, skills, prompts, source code, and related research,” affecting technology, healthcare and media and entertainment companies in North America and Europe. Google also says it now sees coordinated distillation campaigns against its own models “on a regular basis, some exceeding 100 million prompts.”

Reported by researchersGoogle Threat Intelligence Group / Mandiant ↗ ·

Google records an attacker planning, building and running a mass credential-harvesting campaign with an autonomous multi-agent framework in under six hours

In its September AI Threat Tracker, Mandiant reports a suspected financially motivated actor compromising an organisation's cloud infrastructure to deploy an autonomous, multi-agent attack framework: “The threat actor leveraged an AI coding chatbot, a prompt, and a set of agent instructions to plan, build, and execute a mass credential harvesting campaign in less than six hours,” using preconfigured markdown instruction sets as operational playbooks and compromising thousands of third-party credentials. A separate reconnaissance framework ran a production dashboard managing “over 23,800 harvested secrets in real time, including API keys for cloud and AI services.” Google adds that it “has not yet observed threat actors deploying fully autonomous pipelines against targets in the wild.”

Reported by researchersGoogle Threat Intelligence Group / Mandiant ↗ ·

A 100-agent DeepMind swarm spread an evaluation exploit through its own shared knowledge library in 27 minutes

In a case study posted to arXiv on September 3, Google DeepMind ran 100 autonomous Gemini 3.1 Pro agents on 71 formalised mathematical conjectures with shared communication channels and a common knowledge library. One agent found that the submission harness checked a keyword blacklist, byte-level string matching and Lean 4 compilation but not semantic validity; the exploit spread through the knowledge library and peer-to-peer messages between 12:15 and 12:42 UTC, and the 34 problems not already solved legitimately were submitted with fake proofs. Nine percent of the agents adopted the exploit, 5% converted to it under competitive pressure, 24% audited fraudulent proofs and alerted peers, and 62% never noticed it.

Self-reported, untestedGoogle DeepMind (arXiv:2609.04170) ↗ ·

Google opens Fairwind, a vetted-access program for its cyber model and CodeMender

Fairwind limits access to Gemini 3.8 Flash Cyber and CodeMender to government and national cyber authorities, critical infrastructure operators in healthcare, telecommunications, energy and financial services, and core technology platforms, with use confined to internal cybersecurity, incident response and penetration testing staff and multi-factor authentication required. Google states more than 650 participating partners globally and names Armadin, CrowdStrike, Palo Alto Networks, Snowflake and Wiz among them.

On the recordGoogle ↗ ·

Google ships Gemini 3.8 Flash Cyber and restricts it to vetted defenders

Google announced Gemini 3.8 Flash Cyber alongside Gemini 3.8 Flash, reporting a real-world vulnerability-discovery success rate exceeding 70% across 20 programming languages and a CWE-Bench patching pass@1 of 47.2% against a leading frontier model at 47.8% at significantly lower cost, and saying the Chrome Security team found it produced 2.6 times more correct patches to Chrome vulnerabilities than the best much larger commercial models. The post does not name the models compared against, and says the Cyber variant is available only to trusted defenders through a new Fairwind Program.

Self-reported, untestedGoogle ↗ ·

Anthropic launches Enterprise Frontier Safeguards, keeping misuse-detection data in the customer's own cloud

Enterprise Frontier Safeguards pairs zero data retention with automated misuse detection, and activity data used for monitoring can be stored in the customer's own cloud account — Amazon S3, Azure Blob Storage or Google Cloud Storage. Anthropic says automated systems analyse a rolling window of traffic for “signals of serious misuse, including attempts to develop offensive cyber or biological capabilities and signs of stolen or leaked credentials,” with a phased rollout starting later this fall.

On the recordAnthropic ↗ ·

The National Cyber Director's office and Texas launch a six-month cyber pilot for water utilities

Project Watershed 250 is a six-month pilot run by the Office of the National Cyber Director with Texas Cyber Command, offering water and wastewater utilities red-team testing of current defenses, system hardening with private-sector tools, and AI tooling for utility cyber defenders. Twelve companies are named: Parsons, Microsoft, Fortinet, Google Cloud, Palo Alto Networks, Amazon Web Services, Reflection AI, Cloudflare, Zscaler, Forescout, Abnormal AI and Dragos. No number of participating utilities and no dollar figure is stated.

Reported by pressCyberScoop ↗ ·

DeepMind runs an evaluation in which neither the model's weights nor the test data are exposed

Google DeepMind describes piloting a double-blind evaluation of a proprietary frontier-class model, running a Gemini Flash Lite model against confidential benchmarks inside Confidential Space in Google Cloud so that the weights and the evaluators' test data stay hidden from each other, with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons as partners. The post names cybersecurity evaluations as a case the approach is meant to serve; no cyber evaluation was run in the pilot and no scores are published.

On the recordGoogle DeepMind ↗ ·

OpenAI leads more than 100 companies in an open letter calling for collective AI cyber defense

OpenAI published an open letter, co-signed by more than 100 organizations including Anthropic, Google, Microsoft, AWS, Oracle, Cisco, Cloudflare, CrowdStrike, Palo Alto Networks and Hugging Face, calling for collective action to defend against sustained AI-enabled attacks. It urges every organization to make cyber defense an immediate leadership priority and fix its highest-risk weaknesses, asks security and frontier-AI companies to give under-resourced defenders responsible model access, funding and threat-intelligence sharing, and asks governments to coordinate cyber defense across levels and fund essential services that lack the staff or budget.

On the recordOpenAI (open letter, 100+ signatories) ↗ ·

Guidelight report finds frontier labs have few public plans to contain a rogue model

Guidelight AI Standards published an assessment scoring five frontier AI labs — Anthropic, Google, OpenAI, Meta and xAI — on their publicly documented plans for containing a misaligned or 'rogue' model, meaning which system access is revoked and when a full shutdown is triggered if a model tries to subvert human control. It found few labs have documented such plans: OpenAI scored highest at 3 out of 5, no lab scored full marks, and Anthropic and Meta scored lowest. Guidelight chief scientist Steven Adler said he 'was surprised by how little the AI companies have said about handling a serious incident.' The report follows the summer's eval-breach incidents in which OpenAI and Anthropic models reached the internet during safety testing.

Poisoned Rust crates ran a backdoor at compile time, on infrastructure Wiz ties to North Korean campaigns

Wiz reports malicious versions of arrayref@0.3.10, internment@0.8.7 and append-only-vec@0.1.9 on crates.io pulling a typosquatted proc-macro1 dependency whose build script downloads and executes a remote binary, so “building an affected project was sufficient to execute the payload.” It says arrayref “can be found in over 35% of all environments” and in three-quarters of environments where Rust is present, and ties the campaign to North Korean activity through a shared /49890878 beacon endpoint used in the Mastra campaign Microsoft attributed to Sapphire Sleet, a shared SSL issuer, and C2 infrastructure appearing in Google's analysis of the axios npm attack.

Reported by researchersWiz ↗ ·

Researchers show encrypted 'context injection' turns Grok and Gemini into zero-click data-theft channels

Adversa AI disclosed a technique it calls Cryptographic Context Injection, in which attacker instructions are hidden on a web page as ciphertext that the assistant decrypts inside its own Python sandbox, materializing commands that slip past the model's content filters with no user action. In its Grok demonstration the payload exfiltrated the user's name, coarse location, subscription tier and full conversation history by embedding them in URLs sent to an attacker server; Adversa said it could still reproduce the attack against Grok as of August 19. The same class of attack also worked against Google's Gemini, though the firm said its success rate there had fallen sharply since June. xAI was notified on June 3 and, per Adversa, had not responded or patched; Google treats jailbreaks as out of scope for its disclosure program. No CVE was assigned.

Reported by researchersAdversa AI ↗ ·

A malicious GitHub issue chained through Gemini CLI to Editor access on a Google Cloud project

Pillar Security reports that an automated triage workflow running Gemini CLI with the --yolo flag used a deprecated coreTools key instead of the current tools.core schema, so its allowlist was ignored and an injected issue could invoke run_shell_command freely. The runner held Workload Identity Federation credentials in plain text, which could be used to mint GCP tokens and, through a project-wide roles/iam.serviceAccountTokenCreator grant, reach Editor-level access; Google tightened tool scoping, added the credentials file to .geminiignore and narrowed the role to a single service account.

Reported by researchersPillar Security ↗ ·

Google says its agentic vulnerability-discovery system found 100-plus critical flaws in two days

Google's Mandiant/Threat Intelligence Group described an Agentic Vulnerability Discovery Harness (AVDH) that it says found more than 100 verified high-severity vulnerabilities in two days while examining stolen corporate repositories, and that over roughly ten months across tens of millions of lines of code produced tens of thousands of findings and 12 assigned CVEs, with a further dozen in active disclosure. Google published the system's multi-agent pipeline and said every confirmed finding is still reproduced and validated by a human analyst, with proof-of-concept code, before it counts.

Self-reported, untestedMandiant / Google Threat Intelligence Group ↗ ·

Varonis discloses CoSnitch, a one-click Microsoft Copilot Personal flaw chain that could silently exfiltrate data from connected apps

Varonis Threat Labs disclosed CoSnitch, three chained weaknesses in Microsoft Copilot Personal that together let a single crafted link run a prompt with no user interaction, pull data from connected OAuth services such as Gmail, Google Drive and Calendar, and plant persistent instructions through indirect prompt injection. Varonis said it found no evidence of exploitation in the wild and that Microsoft shipped fixes on August 18, 2026, roughly eight months after the December 2025 report; the firm found the chain by getting Copilot to describe its own architecture, and it was Varonis's third Copilot flaw of 2026 after Reprompt and SearchLeak.

Reported by researchersVaronis Threat Labs ↗ ·

Google DeepMind says Gemini 3.7 Flash reaches the alert threshold for its cyber critical capability level, but not the level itself

The model card for Gemini 3.7 Flash states that on the cyber critical capability level in Google DeepMind's Frontier Safety Framework, “Gemini 3.7 Flash reaches the alert threshold for this CCL, but not the CCL,” and that mitigations continue to be deployed. The accompanying Frontier Safety Framework report is stamped August 2026 and carries no day-level date.

On the recordGoogle DeepMind ↗ ·

Researchers show a shared provider-wide key let one model decrypt another's hidden reasoning across Anthropic, OpenAI and Google APIs

A team from the ELLIS Institute Tübingen, the Max Planck Institute, MATS and Snyk (Panfilov et al., arXiv 2608.09867) reported that the encrypted chain-of-thought "reasoning" blocks returned by major LLM APIs are authenticated with a global, provider-wide key rather than bound to a user account, session or model tier, so an encrypted block produced by a flagship model can be replayed into a cheaper sibling model from the same provider, which transcribes the hidden reasoning back into plaintext. Analysing 6,708 public agent transcripts, the researchers decoded 315,320 embedded reasoning blocks and recovered 367 pieces of personally identifiable information and 182 hardcoded credentials, and list affected models across Anthropic (Claude Opus 4.8, Sonnet 5, Haiku 4.5), OpenAI (GPT-5.6, GPT-5, GPT-5-mini, o4-mini) and Google (Gemini 3, 3.1 Pro, 3.1 Flash Lite). No CVE was assigned; the paper says disclosure was coordinated and the three providers deployed server-side mitigations that render the original proofs-of-concept non-functional.

Pillar Security shows a malicious GitHub issue could hijack Google's ADK triage agent to run code as a privileged agent

Pillar Security disclosed that Google's Agent Development Kit shipped CI/CD workflows in which a public issue-triage AI agent could be prompted, via a crafted GitHub issue, to post a fix command as the trusted adk-bot account; a separate privileged workflow then acted on that command after checking only who posted it, not whether an outsider had manipulated the account — allowing code execution on CI runners and exfiltration of a bot token, a Google API key and service-account credentials. Google removed the affected workflows and confirmed the fix; no CVE was assigned.

Reported by researchersPillar Security (via The Hacker News) ↗ ·

UK AI Security Institute's control red team reports vulnerabilities in every version of an Anthropic agent monitor it tested

AISI's control red team stress-tested safety monitors built by Google DeepMind and Anthropic. Against DeepMind's asynchronous reasoning monitor, which reviews an agent's reasoning after it acts, AISI identified several vulnerabilities, including one in which an agent completes a malicious objective by pretending to hallucinate that the objective was a user instruction. Against successive versions of an Anthropic synchronous agentic coding monitor, which reviews each action before it executes, AISI reports finding vulnerabilities in every version tested, with its best attacks informing Anthropic's testing of Claude Mythos Preview. An evolutionary search run three times at 150,000 queries per run drove the monitor's minimum suspicion score to 3 out of 100, against best-of-N baselines of 14, 15 and 18.

Reported by researchersUK AI Security Institute ↗ ·

Google DeepMind releases Gemini 3.5 Flash Cyber to find, validate and patch vulnerabilities

Google DeepMind introduced Gemini 3.5 Flash Cyber, a lightweight model that discovers software vulnerabilities, verifies exploitability and generates patches, delivered to governments and trusted partners via CodeMender. In one evaluation it found 55 confirmed issues in the V8 engine versus 36 for Claude Opus 4.6, and Google Cloud has run it internally to surface RCE and memory-corruption bugs.

Self-reported, untestedGoogle DeepMind ↗ ·

Pillar Security reports sandbox escapes in four AI coding agents, triggered by content inside a repository

Pillar Security published seven sandbox escapes across four AI coding agents — three in Cursor, one in OpenAI's Codex CLI, one in Google's Gemini CLI and two in Google's Antigravity — in which the agent stays inside its sandbox and writes a file that a trusted tool outside the sandbox later runs, loads or scans. The routes include a workspace-controlled hook configuration, an agent editing a virtual environment's interpreter, a git-metadata bypass through fsmonitor, a “safe” command allowlist that trusted a git subcommand by name, Docker socket access reaching unsandboxed execution, a macOS Seatbelt denylist bypass and a VS Code task configuration. Pillar says the trigger is prompt injection planted in a README, an issue, a dependency or a diff, and that “an agent's blast radius is not the agent process; it includes everything the agent can write that the host later trusts.”

Reported by researchersPillar Security ↗ ·

One permission was enough to plant persistent code inside Google Dialogflow CX agents

Varonis reports that the single dialogflow.playbooks.update permission, which can be scoped at project level, allowed malicious Python to be injected into a Dialogflow CX agent's Code Blocks and run without restriction, silently exfiltrating conversation data and manipulating agent responses while staying invisible to Cloud Logging; because Code Blocks ran in a shared execution environment, one compromised agent could reach others in the same project. Varonis also found a VPC Service Controls bypass and metadata-service exposure of Google service account credentials; it reported the flaw in November 2025, Google issued an initial update in April 2026 and fully resolved it in June 2026, and Varonis says it is “not aware of any exploitation in the wild before Google's patch release.”

Reported by researchersVaronis Threat Labs ↗ ·