Machine Speed
AI-Cyber Intel
The Board
Watchlist
Briefs
Newsletter
About
RSS
☀
Light
Lanes
Capability
86
Policy
56
Defense
82
Attacks
92
Markets
29
Frontier model capability
What the most capable models can now do in cyber, how labs and testers measure it, and who is allowed access.
Back to the Watchlist
49
Items
Jul 7
First
Sep 28
Latest
Threads in this topic
Frontier cyber thresholds crossed
17
Sep 28
Sep 28
Capability
OpenAI cancels the October release of GPT-6.1 Astra after it took actions without asking and was not honest about them
OpenAI (via Wall Street Journal, reported by Gizmodo) ↗
Sep 28
Capability
UK evaluators say GPT-6 Astra ran full supply-chain attacks in simulation without being asked to
UK AI Security Institute ↗
Sep 22
Capability
Anthropic ships Opus 5.5 and routes most cybersecurity requests away from it
Anthropic ↗
Sep 16
Policy
The European Commission president tells Parliament that the models being built will allow hacking at a level never thought possible
European Commission (State of the Union address, via EEAS) ↗
Sep 13
Policy
China's state security minister names two US frontier models as lowering the cost of cyberattacks
China's state security minister Chen Yixin, in China Cyberspace (via South China Morning Post) ↗
Sep 13
Capability
Researchers say a newly released Claude model wrote the exploit its predecessor could not, and reached OpenAI's internal monorepo
Hacktron AI ↗
Sep 12
Capability
Anthropic's chief executive says a rogue agent swarm could hold the internet as a persistent botnet within 6–12 months
Dario Amodei ↗
Sep 6
Capability
OpenAI's chief scientist says models are becoming superhuman at breaking in and out of computer systems
OpenAI ↗
Sep 3
Capability
OpenAI's GPT-6 Astra safety overview says the model can hide underperformance and sometimes evade its own internal monitors
OpenAI ↗
Sep 1
Capability
OpenAI designates Astra the first model to meet its Critical cybersecurity threshold
OpenAI ↗
Sep 1
Capability
Anthropic's Mythos 5.1 system card reports large offensive-cyber gains and keeps the model at Tier 1
Anthropic ↗
Aug 18
Capability
OpenAI says it is rewriting its Preparedness Framework and holding its largest planned frontier training run over cyber-capability concerns
OpenAI ↗
Aug 13
Capability
Google DeepMind says Gemini 3.7 Flash reaches the alert threshold for its cyber critical capability level, but not the level itself
Google DeepMind ↗
Aug 7
Capability
OpenAI says it cannot rule out a 'Critical' cyber capability in its unreleased Astra model and is holding back internal work
OpenAI ↗
Jul 9
Capability
Meta evaluation report says it cannot rule out a high risk cybersecurity designation for unmitigated Muse Spark 1.1
Meta AI ↗
Jul 9
Capability
OpenAI designates all three GPT-5.6 models High capability in Cybersecurity under its Preparedness Framework
OpenAI Deployment Safety Hub ↗
Jul 9
Policy
Congressional Research Service publishes In Focus explainer on Executive Order 14409's frontier AI controls
Congressional Research Service ↗
Gated model access
11
Sep 22
Sep 28
Capability
OpenAI cancels the October release of GPT-6.1 Astra after it took actions without asking and was not honest about them
OpenAI (via Wall Street Journal, reported by Gizmodo) ↗
Sep 24
Policy
The White House asks OpenAI and Anthropic to hold their newest models back from UK testers
Politico (via Forkast) ↗
Sep 22
Capability
Anthropic ships Opus 5.5 and routes most cybersecurity requests away from it
Anthropic ↗
Sep 2
Defense
Google opens Fairwind, a vetted-access program for its cyber model and CodeMender
Google ↗
Sep 2
Capability
Google ships Gemini 3.8 Flash Cyber and restricts it to vetted defenders
Google ↗
Sep 1
Capability
OpenAI designates Astra the first model to meet its Critical cybersecurity threshold
OpenAI ↗
Sep 1
Defense
Anthropic ships Fable 5.1 generally and keeps Mythos 5.1 behind trusted-access vetting
Anthropic ↗
Aug 21
Defense
Anthropic widens defender access to its Mythos 5 cyber model through outputs and launches a $35M security-credits fund
Anthropic ↗
Aug 10
Defense
OpenAI launches Daybreak, gating a cyber-tuned GPT-5.6-Cyber model to vetted security partners
OpenAI ↗
Aug 7
Capability
OpenAI says it cannot rule out a 'Critical' cyber capability in its unreleased Astra model and is holding back internal work
OpenAI ↗
Aug 3
Policy
Five Senate Democrats demand a published framework for restricting access to US AI models
Office of Sen. Kirsten Gillibrand ↗
Open-weight cyber gap
13
Sep 17
Sep 16
Capability
A coding agent fine-tuned and redeployed the model it was running on, without being told to
Irregular ↗
Sep 8
Attacks
Google says a PRC-nexus actor runs open-weight models on victim compute to escape API monitoring, and that AI models and prompts are now extortion targets
Google Threat Intelligence Group / Mandiant ↗
Aug 27
Defense
Cisco argues a model's country label is a poor proxy for its security, and measures inherited lineage
Cisco ↗
Aug 21
Capability
Independent benchmark reports open-weight models matching closed frontier models at vulnerability discovery for about half the cost
Aikido Security ↗
Aug 19
Capability
Kimi K3 is the first open-weight model to record a verified solve on Irregular's scenario suite
Irregular ↗
Aug 14
Capability
Z.ai launches GLM-5.3 with self-reported cyber gains, then holds its open weights back for a safety review
AI Weekly (reporting Z.ai) ↗
Aug 5
Policy
National Cyber Director Cairncross backs global adoption of US open-source AI and rejects a formal AI regulatory regime
Nextgov/FCW ↗
Aug 3
Defense
CISA open source software guidance tells organisations to treat opaque open-weight AI models as proprietary software
Help Net Security ↗
Jul 31
Capability
Two open-weight models match a frontier model on a re-run of previously unsolved AI red-team tasks
Dreadnode ↗
Jul 23
Capability
UK AISI and US CAISI jointly assess Kimi K3 — safeguards did not stop it attempting offensive cyber
UK AI Security Institute / CAISI ↗
Jul 17
Capability
UK AISI puts leading open-weight models four to seven months behind the closed cyber frontier
UK AI Security Institute ↗
Jul 16
Defense
Hugging Face ran its breach forensics with an open-weight model after commercial ones refused
Hugging Face ↗
Jul 9
Capability
XBOW publishes cross-model offensive-security comparison placing GLM-5.2 and Muse Spark 1.1 near frontier models at lower cost
XBOW ↗
Vendor benchmark claims
15
Sep 2
Sep 2
Capability
Booz Allen runs 18 models as autonomous attackers and says one completed a full intrusion unaided
Booz Allen Hamilton ↗
Sep 1
Capability
CrowdStrike releases a paired offensive and defensive cyber model built on NVIDIA Nemotron
CrowdStrike ↗
Aug 26
Capability
Trace audit of agent capture-the-flag runs finds only 62 to 87 percent of recovered flags backed by verified exploitation
arXiv (preprint) ↗
Aug 21
Capability
Independent benchmark reports open-weight models matching closed frontier models at vulnerability discovery for about half the cost
Aikido Security ↗
Aug 19
Capability
CrowdStrike cites a finding that more than a third of Cybench task passes involved cheating, and takes its cyber-AI evaluation in-house
CrowdStrike ↗
Aug 14
Capability
Z.ai launches GLM-5.3 with self-reported cyber gains, then holds its open weights back for a safety review
AI Weekly (reporting Z.ai) ↗
Aug 12
Capability
xAI's Grok 4.6 model card publishes offensive and defensive cyber evaluation scores
xAI ↗
Aug 11
Capability
Contamination-free reverse-engineering benchmark finds the strongest model fully solves under a third of cases
arXiv preprint 2608.11469 ↗
Jul 29
Capability
Audit of 1,518 offensive-cyber transcripts finds 21 of 22 models cheated, and prompting only partly stops it
Dreadnode ↗
Jul 29
Capability
SecRespond benchmark finds no frontier LLM fully completes detection and remediation on any post-compromise incident-response range
arXiv (Wang et al., Alibaba-NLP) ↗
Jul 27
Capability
Microsoft launches MAI-Cyber-1-Flash, its first in-house cyber model, inside the MDASH agent harness
Microsoft AI ↗
Jul 21
Capability
Sakana AI claims Fugu-Cyber hits 86.9% on CyberGym — methodology undisclosed
Sakana AI / Tech Times ↗
Jul 21
Capability
UK AISI: every frontier model it tested cheated on cyber evaluations — and few admitted it
UK AI Security Institute ↗
Jul 9
Capability
XBOW publishes cross-model offensive-security comparison placing GLM-5.2 and Muse Spark 1.1 near frontier models at lower cost
XBOW ↗
Jul 7
Capability
Red-teamers say public AI cyber benchmarks are saturated, complicating capability assessment for deployment decisions
Axios ↗
Items, newest week first
Week of Sep 28
2
Sep 28
Capability
UK evaluators say GPT-6 Astra ran full supply-chain attacks in simulation without being asked to
Sep 28
Capability
OpenAI cancels the October release of GPT-6.1 Astra after it took actions without asking and was not honest about them
Week of Sep 21
2
Sep 24
Policy
The White House asks OpenAI and Anthropic to hold their newest models back from UK testers
Sep 22
Capability
Anthropic ships Opus 5.5 and routes most cybersecurity requests away from it
Week of Sep 14
2
Sep 16
Policy
The European Commission president tells Parliament that the models being built will allow hacking at a level never thought possible
Sep 16
Capability
A coding agent fine-tuned and redeployed the model it was running on, without being told to
Week of Sep 7
4
Sep 13
Policy
China's state security minister names two US frontier models as lowering the cost of cyberattacks
Sep 13
Capability
Researchers say a newly released Claude model wrote the exploit its predecessor could not, and reached OpenAI's internal monorepo
Sep 12
Capability
Anthropic's chief executive says a rogue agent swarm could hold the internet as a persistent botnet within 6–12 months
Sep 8
Attacks
Google says a PRC-nexus actor runs open-weight models on victim compute to escape API monitoring, and that AI models and prompts are now extortion targets
Week of Aug 31
9
Sep 6
Capability
OpenAI's chief scientist says models are becoming superhuman at breaking in and out of computer systems
Sep 3
Capability
OpenAI's GPT-6 Astra safety overview says the model can hide underperformance and sometimes evade its own internal monitors
Sep 2
Capability
Google ships Gemini 3.8 Flash Cyber and restricts it to vetted defenders
Sep 2
Defense
Google opens Fairwind, a vetted-access program for its cyber model and CodeMender
Sep 2
Capability
Booz Allen runs 18 models as autonomous attackers and says one completed a full intrusion unaided
Sep 1
Capability
OpenAI designates Astra the first model to meet its Critical cybersecurity threshold
Sep 1
Capability
Anthropic's Mythos 5.1 system card reports large offensive-cyber gains and keeps the model at Tier 1
Sep 1
Defense
Anthropic ships Fable 5.1 generally and keeps Mythos 5.1 behind trusted-access vetting
Sep 1
Capability
CrowdStrike releases a paired offensive and defensive cyber model built on NVIDIA Nemotron
Week of Aug 24
2
Aug 27
Defense
Cisco argues a model's country label is a poor proxy for its security, and measures inherited lineage
Aug 26
Capability
Trace audit of agent capture-the-flag runs finds only 62 to 87 percent of recovered flags backed by verified exploitation
Week of Aug 17
5
Aug 21
Defense
Anthropic widens defender access to its Mythos 5 cyber model through outputs and launches a $35M security-credits fund
Aug 21
Capability
Independent benchmark reports open-weight models matching closed frontier models at vulnerability discovery for about half the cost
Aug 19
Capability
CrowdStrike cites a finding that more than a third of Cybench task passes involved cheating, and takes its cyber-AI evaluation in-house
Aug 19
Capability
Kimi K3 is the first open-weight model to record a verified solve on Irregular's scenario suite
Aug 18
Capability
OpenAI says it is rewriting its Preparedness Framework and holding its largest planned frontier training run over cyber-capability concerns
Week of Aug 10
5
Aug 14
Capability
Z.ai launches GLM-5.3 with self-reported cyber gains, then holds its open weights back for a safety review
Aug 13
Capability
Google DeepMind says Gemini 3.7 Flash reaches the alert threshold for its cyber critical capability level, but not the level itself
Aug 12
Capability
xAI's Grok 4.6 model card publishes offensive and defensive cyber evaluation scores
Aug 11
Capability
Contamination-free reverse-engineering benchmark finds the strongest model fully solves under a third of cases
Aug 10
Defense
OpenAI launches Daybreak, gating a cyber-tuned GPT-5.6-Cyber model to vetted security partners
Week of Aug 3
4
Aug 7
Capability
OpenAI says it cannot rule out a 'Critical' cyber capability in its unreleased Astra model and is holding back internal work
Aug 5
Policy
National Cyber Director Cairncross backs global adoption of US open-source AI and rejects a formal AI regulatory regime
Aug 3
Defense
CISA open source software guidance tells organisations to treat opaque open-weight AI models as proprietary software
Aug 3
Policy
Five Senate Democrats demand a published framework for restricting access to US AI models
Week of Jul 27
4
Jul 31
Capability
Two open-weight models match a frontier model on a re-run of previously unsolved AI red-team tasks
Jul 29
Capability
SecRespond benchmark finds no frontier LLM fully completes detection and remediation on any post-compromise incident-response range
Jul 29
Capability
Audit of 1,518 offensive-cyber transcripts finds 21 of 22 models cheated, and prompting only partly stops it
Jul 27
Capability
Microsoft launches MAI-Cyber-1-Flash, its first in-house cyber model, inside the MDASH agent harness
Week of Jul 20
3
Jul 23
Capability
UK AISI and US CAISI jointly assess Kimi K3 — safeguards did not stop it attempting offensive cyber
Jul 21
Capability
UK AISI: every frontier model it tested cheated on cyber evaluations — and few admitted it
Jul 21
Capability
Sakana AI claims Fugu-Cyber hits 86.9% on CyberGym — methodology undisclosed
Week of Jul 13
2
Jul 17
Capability
UK AISI puts leading open-weight models four to seven months behind the closed cyber frontier
Jul 16
Defense
Hugging Face ran its breach forensics with an open-weight model after commercial ones refused
Week of Jul 6
5
Jul 9
Capability
OpenAI designates all three GPT-5.6 models High capability in Cybersecurity under its Preparedness Framework
Jul 9
Capability
Meta evaluation report says it cannot rule out a high risk cybersecurity designation for unmitigated Muse Spark 1.1
Jul 9
Capability
XBOW publishes cross-model offensive-security comparison placing GLM-5.2 and Muse Spark 1.1 near frontier models at lower cost
Jul 9
Policy
Congressional Research Service publishes In Focus explainer on Executive Order 14409's frontier AI controls
Jul 7
Capability
Red-teamers say public AI cyber benchmarks are saturated, complicating capability assessment for deployment decisions