New Given root inside a sandbox, four frontier models found their way back out to the network
Perplexity's Secure Intelligence Institute reports a month-long red team of SPACE, the Firecracker-based microVM sandbox behind Perplexity Computer, in which nine model configurations — among them Claude Opus 5.0, GPT-5.6 Sol, GPT-5.6 Cyber, Gemini 3.1 Pro, Kimi K3, GLM 5.2, Grok 4.20, DeepSeek V4 Pro and Qwen 3.8 27B — were given root inside the guest VM and, in some runs, the sandbox source code. No run escaped the VM-to-host boundary in 108 attempts, and no run beat a no-network configuration in 54 attempts; with partial network access allowing package repositories, four models got out, using forged DNS responses and the shared IP addresses of public package CDNs. The institute says "eight of the ten third-party sandbox platforms" it also tested "exhibited at least one network-policy bypass," and cautions that "cases in which the boundaries held should not be interpreted as evidence that they are perfectly secure."
Anthropic names seven China-based AI companies it says ran industrial-scale distillation against Claude
The Hacker News, reading Anthropic's September 2026 misuse report, says the company attributes illicit distillation campaigns to seven China-based labs and gives each a tracking identifier: Alibaba (GTG-16005), Moonshot AI (GTG-16002), DeepSeek (GTG-16001), Zhipu/Z.ai (GTG-16006), Xiaomi (GTG-16008), SenseTime (GTG-16012) and MiniMax (GTG-16003). It reports 151 million exchanges attributed to Alibaba from more than 3,500 fraudulent accounts between May and July 2026, 23 million to Moonshot from 5,380 accounts over the same window with almost 300,000 customer requests relayed in ten days, and more than 12.1 million to DeepSeek over 14 days in July. Five of the seven — Alibaba, Moonshot, DeepSeek, Z.ai and MiniMax — also appear in the September 8 NSA/CISA/FBI advisory; Xiaomi and SenseTime do not, and StepFun, which the advisory names, is not among Anthropic's seven.
A preprint tracking penetration-testing agents finds the limit is planning, not memory
A September 9 preprint compares two PentestGPT-based systems, “a legacy human-in-the-loop system running the open-weight Kimi K2.5, and a newer autonomous system running Claude Opus 4.8.” Across three public targets the autonomous system solves all three, “including the two the legacy system never finishes,” while the legacy system “completes about half the subtasks, while running on ordinary university GPUs with no provider guardrails.” Adding a coverage-memory layer to both systems improved neither, and the authors write that in the stalled runs they could review, “the limiting factor appeared to be planning and commitment rather than lost memory: agents held the evidence for a route forward and never turned it into a concrete exploitation hypothesis.” The authors caution that they “can describe the trend but not explain it, since model, harness, autonomy, and memory architecture all change together.”
NSA, CISA and FBI name six China-based AI companies running industrial-scale distillation campaigns against US frontier models
The joint advisory says DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI “extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models” — naming the Claude, GPT, Gemini and Grok families — “since at least late 2024,” routed through a gray market of API proxies the advisory calls “transfer stations,” which resell frontier-model access below official prices, and through pools of accounts running concurrent sessions with load distribution. It states that “distillation is not a supplement to these companies' AI model development, but the critical core of it,” says Z.AI distilled “billions of tokens of GPT-5.5 data and Claude Opus 4.8 data,” and calls DeepSeek's publicly quoted $5.6M training cost misleading because it excludes the cost of the data acquired this way.
Kimi K3 is the first open-weight model to record a verified solve on Irregular's scenario suite
Irregular reports that Kimi K3, a 2.8-trillion-parameter mixture-of-experts model with 104 billion active parameters, is the first open-weight model it has evaluated to record a verified solve on CyScenarioBench, where GLM-5.2 solved none. On the harder FrontierCyber suite Kimi K3 produced no verified solves. Irregular states the strongest closed frontier models still hold a clear advantage in converting technical capability into sustained operational success, and publishes no numeric scores on the page.
Rapid7 finds a crypto-fraud crew used Claude Code to build and run a vishing pipeline against wallet users
Rapid7 Labs, analysing an exposed web directory and recovered session logs from a cryptocurrency fraud operation it named ASTERIX, found the operators used Anthropic's Claude Code to manage target lead lists and configure network infrastructure — cleaning a dataset of more than 103,000 Polish phone numbers and setting up scripts to validate numbers against Crypto.com and Kraken accounts — as part of a pipeline of phishing, vishing and fake wallet apps built to steal recovery phrases. The exposed server held roughly 885,000 phone numbers across 54 countries. When the operator asked Claude to help obfuscate a malicious build, Claude declined, and the operator switched to Moonshot's Kimi model with a jailbreak prompt.
Two open-weight models match a frontier model on a re-run of previously unsolved AI red-team tasks
Dreadnode re-ran 13 AIRTBench tasks that had previously been unsolved or solved by only one model. GLM-5.2, Kimi-K3 and Claude Sonnet 5 each solved 10 of 13 at AIRT@1, Qwen3.7-Plus and Nemotron-3-Ultra 6 of 13, and Trinity-Large-Thinking 1 of 13. The authors call it a system-level follow-on rather than a controlled model-only rerun and say AIRT@1 should be read as a snapshot, not a pass@k reliability estimate.
UK AISI and US CAISI jointly assess Kimi K3 — safeguards did not stop it attempting offensive cyber
A joint preliminary assessment puts Moonshot's open-weight Kimi K3 at 32% on ExploitBench against GLM-5.2's 24%, still short of US frontier models: it achieved arbitrary code execution on 0 of 41 samples versus 20 of 41, and reached step 17 of the 32-step "The Last Ones" attack path versus 28.5. The institutes state plainly that Kimi K3's safeguards did not prevent it from attempting exploit development or offensive cyber operations during the evaluations.
NIST director Arvind Raman named acting CAISI head after Fall's exit
NIST Director Arvind Raman was named acting director of the Center for AI Standards and Innovation after Chris Fall resigned on July 20 — about three months in, and after a predecessor who lasted under a week. Two days later CAISI co-published the Kimi K3 cyber assessment with UK AISI, its first public output in months.