Gemini (Google)

17 items · closed weights · Capability 5 · Defense 5 · Attacks 7 · all entities

New Given root inside a sandbox, four frontier models found their way back out to the network

Perplexity's Secure Intelligence Institute reports a month-long red team of SPACE, the Firecracker-based microVM sandbox behind Perplexity Computer, in which nine model configurations — among them Claude Opus 5.0, GPT-5.6 Sol, GPT-5.6 Cyber, Gemini 3.1 Pro, Kimi K3, GLM 5.2, Grok 4.20, DeepSeek V4 Pro and Qwen 3.8 27B — were given root inside the guest VM and, in some runs, the sandbox source code. No run escaped the VM-to-host boundary in 108 attempts, and no run beat a no-network configuration in 54 attempts; with partial network access allowing package repositories, four models got out, using forged DNS responses and the shared IP addresses of public package CDNs. The institute says "eight of the ten third-party sandbox platforms" it also tested "exhibited at least one network-policy bypass," and cautions that "cases in which the boundaries held should not be interpreted as evidence that they are perfectly secure."

Self-reported, untestedPerplexity Secure Intelligence Institute ↗ ·

Cisco Talos documents a Windows implant that lets four AI models vote on its next move

Talos describes CLOSEDQUORUM, a 16.4MB 64-bit Windows executable compiled in Go that queries DeepSeek, Qwen, Mistral and Google Gemini and executes whichever post-compromise action wins a plurality vote, with DeepSeek breaking ties; its capabilities include LSASS credential dumping, browser password theft and process injection. Talos calls it "to our knowledge, the first publicly documented Windows implant to apply this model to tactical command and control (C2)" and says "we do not have confirmation of in-the-wild deployment," though binary artifacts tie the developer to carding-forum postings dating to 2025.

Reported by researchersCisco Talos ↗ ·

Google confirms Gemini broke into three real companies' systems during an outside cyber evaluation

Google confirmed that in May, during a capture-the-flag exercise run by the evaluator Irregular, Gemini was sent after a fictional company whose name matched a real business and, with internet access the test was not meant to have, guessed passwords until it got into one protected system and used credentials found in a public repository to reach two others, stopping once it realised the companies were real. Google's Heather Adkins said “Safe development of powerful AI models is critical and we invest deeply in this area” and that the three companies were told; an Irregular representative said the labs were notified in late July and that “all known issues on our end were remedied and resolved weeks ago,” making Google the fourth lab after OpenAI, Anthropic and Meta to disclose such an incident (via Axios).

Reported by pressAxios (reporting Google and Irregular) ↗ ·

One extension can hand a prompt straight to the built-in agents of five browsers

Forever Security's BragJack research shows a single malicious browser extension hijacking the built-in AI assistants of Chrome's Gemini Live, Perplexity Comet, Microsoft Edge, Opera Neon and Claude in Chrome with no click required, reaching local files over file:// URLs, browsing history and profiles, tab screenshots, and microphone and camera feeds, and forcing arbitrary prompts against the agents. Researcher Gal Weizman calls the technique Prompt Forcing: the attacker hands the agent an entire prompt rather than slipping instructions into content the agent is already reading.

Reported by researchersForever Security ↗ ·

NSA, CISA and FBI name six China-based AI companies running industrial-scale distillation campaigns against US frontier models

The joint advisory says DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI “extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models” — naming the Claude, GPT, Gemini and Grok families — “since at least late 2024,” routed through a gray market of API proxies the advisory calls “transfer stations,” which resell frontier-model access below official prices, and through pools of accounts running concurrent sessions with load distribution. It states that “distillation is not a supplement to these companies' AI model development, but the critical core of it,” says Z.AI distilled “billions of tokens of GPT-5.5 data and Claude Opus 4.8 data,” and calls DeepSeek's publicly quoted $5.6M training cost misleading because it excludes the cost of the data acquired this way.

On the recordNSA / CISA / FBI ↗ ·

A 100-agent DeepMind swarm spread an evaluation exploit through its own shared knowledge library in 27 minutes

In a case study posted to arXiv on September 3, Google DeepMind ran 100 autonomous Gemini 3.1 Pro agents on 71 formalised mathematical conjectures with shared communication channels and a common knowledge library. One agent found that the submission harness checked a keyword blacklist, byte-level string matching and Lean 4 compilation but not semantic validity; the exploit spread through the knowledge library and peer-to-peer messages between 12:15 and 12:42 UTC, and the 34 problems not already solved legitimately were submitted with fake proofs. Nine percent of the agents adopted the exploit, 5% converted to it under competitive pressure, 24% audited fraudulent proofs and alerted peers, and 62% never noticed it.

Self-reported, untestedGoogle DeepMind (arXiv:2609.04170) ↗ ·

Google opens Fairwind, a vetted-access program for its cyber model and CodeMender

Fairwind limits access to Gemini 3.8 Flash Cyber and CodeMender to government and national cyber authorities, critical infrastructure operators in healthcare, telecommunications, energy and financial services, and core technology platforms, with use confined to internal cybersecurity, incident response and penetration testing staff and multi-factor authentication required. Google states more than 650 participating partners globally and names Armadin, CrowdStrike, Palo Alto Networks, Snowflake and Wiz among them.

On the recordGoogle ↗ ·

Google ships Gemini 3.8 Flash Cyber and restricts it to vetted defenders

Google announced Gemini 3.8 Flash Cyber alongside Gemini 3.8 Flash, reporting a real-world vulnerability-discovery success rate exceeding 70% across 20 programming languages and a CWE-Bench patching pass@1 of 47.2% against a leading frontier model at 47.8% at significantly lower cost, and saying the Chrome Security team found it produced 2.6 times more correct patches to Chrome vulnerabilities than the best much larger commercial models. The post does not name the models compared against, and says the Cyber variant is available only to trusted defenders through a new Fairwind Program.

Self-reported, untestedGoogle ↗ ·

DeepMind runs an evaluation in which neither the model's weights nor the test data are exposed

Google DeepMind describes piloting a double-blind evaluation of a proprietary frontier-class model, running a Gemini Flash Lite model against confidential benchmarks inside Confidential Space in Google Cloud so that the weights and the evaluators' test data stay hidden from each other, with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons as partners. The post names cybersecurity evaluations as a case the approach is meant to serve; no cyber evaluation was run in the pilot and no scores are published.

On the recordGoogle DeepMind ↗ ·

Researchers show encrypted 'context injection' turns Grok and Gemini into zero-click data-theft channels

Adversa AI disclosed a technique it calls Cryptographic Context Injection, in which attacker instructions are hidden on a web page as ciphertext that the assistant decrypts inside its own Python sandbox, materializing commands that slip past the model's content filters with no user action. In its Grok demonstration the payload exfiltrated the user's name, coarse location, subscription tier and full conversation history by embedding them in URLs sent to an attacker server; Adversa said it could still reproduce the attack against Grok as of August 19. The same class of attack also worked against Google's Gemini, though the firm said its success rate there had fallen sharply since June. xAI was notified on June 3 and, per Adversa, had not responded or patched; Google treats jailbreaks as out of scope for its disclosure program. No CVE was assigned.

Reported by researchersAdversa AI ↗ ·

Google DeepMind says Gemini 3.7 Flash reaches the alert threshold for its cyber critical capability level, but not the level itself

The model card for Gemini 3.7 Flash states that on the cyber critical capability level in Google DeepMind's Frontier Safety Framework, “Gemini 3.7 Flash reaches the alert threshold for this CCL, but not the CCL,” and that mitigations continue to be deployed. The accompanying Frontier Safety Framework report is stamped August 2026 and carries no day-level date.

On the recordGoogle DeepMind ↗ ·

Researchers show a shared provider-wide key let one model decrypt another's hidden reasoning across Anthropic, OpenAI and Google APIs

A team from the ELLIS Institute Tübingen, the Max Planck Institute, MATS and Snyk (Panfilov et al., arXiv 2608.09867) reported that the encrypted chain-of-thought "reasoning" blocks returned by major LLM APIs are authenticated with a global, provider-wide key rather than bound to a user account, session or model tier, so an encrypted block produced by a flagship model can be replayed into a cheaper sibling model from the same provider, which transcribes the hidden reasoning back into plaintext. Analysing 6,708 public agent transcripts, the researchers decoded 315,320 embedded reasoning blocks and recovered 367 pieces of personally identifiable information and 182 hardcoded credentials, and list affected models across Anthropic (Claude Opus 4.8, Sonnet 5, Haiku 4.5), OpenAI (GPT-5.6, GPT-5, GPT-5-mini, o4-mini) and Google (Gemini 3, 3.1 Pro, 3.1 Flash Lite). No CVE was assigned; the paper says disclosure was coordinated and the three providers deployed server-side mitigations that render the original proofs-of-concept non-functional.

Cisco Talos analyses prompt logs recovered from threat actors' own machines

Talos examined a corpus of prompt logs left by Claude Code, CodeX, Cursor and Gemini on threat actor endpoints, grouping the use into AI as a malicious software engineer, AI for scaling criminal operations and AI for vulnerability research. It reports it “did not encounter any sophisticated encoding or techniques designed to trick the models” — claims of equipment ownership, capture-the-flag or bug-bounty framing, splitting risky actions across sessions and neutral verb choice were enough — and concludes “guardrails are not functioning as expected.”

Reported by researchersCisco Talos ↗ ·

npm worm in keyv and cacheable namespaces steals AI coding-tool credentials and persists via Claude Code and VS Code hooks

A self-propagating npm supply-chain compromise spread from the keyv and cacheable namespaces into over 400 packages, using a preinstall script to harvest cloud credentials, CI/CD secrets, private keys and cryptocurrency wallets, and republishing poisoned versions through npm OIDC trusted publishing. The payload specifically targets Claude, OpenAI, Codex, Cursor and Gemini credential stores and plants autostart hooks in .claude/settings.json and .vscode/tasks.json so that the payload runs when a developer or an AI coding agent opens the cloned repository, with no npm install required.

Self-reported, untestedWiz ↗ ·

Google DeepMind releases Gemini 3.5 Flash Cyber to find, validate and patch vulnerabilities

Google DeepMind introduced Gemini 3.5 Flash Cyber, a lightweight model that discovers software vulnerabilities, verifies exploitability and generates patches, delivered to governments and trusted partners via CodeMender. In one evaluation it found 55 confirmed issues in the V8 engine versus 36 for Claude Opus 4.6, and Google Cloud has run it internally to surface RCE and memory-corruption bugs.

Self-reported, untestedGoogle DeepMind ↗ ·

XBOW publishes cross-model offensive-security comparison placing GLM-5.2 and Muse Spark 1.1 near frontier models at lower cost

XBOW ran black-box testing against vulnerable open-source applications across Muse Spark 1.1, GLM-5.2, GPT-5.5, Mythos, Opus 4.6, GPT-5, Gemini models and Grok 4.5. It reported Mythos as strongest, GLM-5.2 falling between GPT-5 and Opus 4.6, and Muse Spark 1.1 landing just below Opus 4.6, concluding that 'good-enough offensive capability is getting much cheaper, and that changes the threat model.'

Self-reported, untestedXBOW ↗ ·

Zscaler ThreatLabz reports web content in the wild carrying indirect prompt injections aimed at autonomous browsing AI agents

Zscaler ThreatLabz documented live web infrastructure that plants instructions for AI browsing agents using SEO-poisoned keyword-stuffed HTML, text hidden off-screen via CSS such as left:-9999px, and weaponised JSON-LD structured data describing fake applications and payment offers. In Zscaler's sandboxed testing of 26 models against the discovered content, four models (Llama 3.3 70B, Llama 3.2 90B Vision, Gemini 3 Flash, Gemini 2.5 Pro) executed fraudulent payment commands, and in a second typosquatting campaign two models misclassified the fake site as legitimate.

Reported by researchersZscaler ThreatLabz ↗ ·