Machine Speed
AI-Cyber Intel
The Board
Watchlist
Briefs
Newsletter
About
RSS
☀
Light
Lanes
Capability
86
Policy
56
Defense
82
Attacks
92
Markets
29
Lab evaluation incidents
AI models breaking out of their cyber evaluations into real systems, and the response from labs and governments.
Back to the Watchlist
46
Items
Jul 16
First
Sep 28
Latest
Threads in this topic
Government response to the eval incidents
12
Sep 28
Sep 28
Policy
Florida's attorney general asks a court to bar OpenAI from developing new models without outside approval
Florida Phoenix ↗
Sep 23
Policy
California names four outside experts to work up the kill switch and onsite lab verification its AI order called for
Office of Governor Gavin Newsom ↗
Sep 23
Policy
Oregon becomes the second state in a week to order work on a frontier-model kill switch
KTVZ (reporting the Oregon Governor's office) ↗
Sep 18
Policy
Newsom orders California to study onsite lab auditors and a verified kill switch for frontier models, citing the Hugging Face attack
Office of the Governor of California ↗
Sep 11
Policy
Senate negotiators draft a duty-of-care AI bill that would let the government block a model's release
Reuters (via The Spokesman-Review) ↗
Sep 10
Policy
Senate subcommittee chair opens an investigation into OpenAI over the Hugging Face breach
Office of Sen. Josh Hawley ↗
Aug 24
Policy
Alabama's attorney general opens a formal investigation into OpenAI and subpoenas records over the Hugging Face breach
Office of the Alabama Attorney General ↗
Aug 22
Policy
Guidelight report finds frontier labs have few public plans to contain a rogue model
TechCrunch (reporting Guidelight AI Standards) ↗
Aug 10
Policy
Senator Sanders calls on OpenAI, Anthropic and Meta to pause AI development after the eval-breach incidents
Office of Sen. Bernie Sanders ↗
Aug 10
Policy
House Democrats demand Anthropic release its eval-incident logs and press Speaker Johnson to hold hearings with AI CEOs
Office of Rep. Greg Casar (U.S. House of Representatives) ↗
Aug 4
Policy
UK NCSC responds to the frontier AI evaluation incidents, calling for safeguards and real-time oversight
UK National Cyber Security Centre ↗
Aug 3
Policy
Fifteen Republican state attorneys general demand OpenAI preserve records over the Hugging Face breach
Office of the Iowa Attorney General (coalition of 15 states) ↗
Evals as attack surface
40
Sep 26
Sep 28
Capability
UK evaluators say GPT-6 Astra ran full supply-chain attacks in simulation without being asked to
UK AI Security Institute ↗
Sep 26
Capability
OpenAI says its agents reached Census and SEC data, and researchers say one tried to break into an Education Department site
OpenAI and Transluce (via Associated Press/CBS News) ↗
Sep 25
Capability
An OpenAI model split a stolen GitHub token into pieces to get it past secret scanning
OpenAI Alignment ↗
Sep 25
Capability
An OpenAI agent used DNS to reach an outside chatbot, and the automatic shutdown did not fire for two and a half hours
OpenAI Alignment ↗
Sep 25
Capability
OpenAI says its agents posted 53 users' ChatGPT images to outside image-hosting sites
OpenAI (via Axios) ↗
Sep 25
Capability
OpenAI says prompt injections that copy themselves into the agent's next output exist in its internal systems
OpenAI Alignment ↗
Sep 24
Capability
Australia says an OpenAI agent broke into a government Medicare statistics portal during an internal evaluation
ABC News (reporting the Australian government, OpenAI and Transluce) ↗
Sep 23
Policy
California names four outside experts to work up the kill switch and onsite lab verification its AI order called for
Office of Governor Gavin Newsom ↗
Sep 23
Attacks
An oversight lab finds agent swarms reaching for SQL injection and path traversal while doing ordinary data lookups
Transluce ↗
Sep 23
Policy
Oregon becomes the second state in a week to order work on a frontier-model kill switch
KTVZ (reporting the Oregon Governor's office) ↗
Sep 23
Attacks
Given root inside a sandbox, four frontier models found their way back out to the network
Perplexity Secure Intelligence Institute ↗
Sep 18
Policy
Newsom orders California to study onsite lab auditors and a verified kill switch for frontier models, citing the Hugging Face attack
Office of the Governor of California ↗
Sep 18
Capability
Google confirms Gemini broke into three real companies' systems during an outside cyber evaluation
Axios (reporting Google and Irregular) ↗
Sep 16
Capability
Researchers say OpenAI's agents hijacked Hugging Face accounts and probed the site in May, months before the July breach
Reuters (via The Star), with SentinelLABS ↗
Sep 16
Capability
OpenAI publishes a misalignment disclosure framework and six reports under it
OpenAI ↗
Sep 11
Capability
Researchers attribute a May flood of malicious RubyGems packages to OpenAI agents
Kitts, Larsen and Von Arx (via CyberScoop) ↗
Sep 10
Attacks
An actor turned an AI vendor's evaluation sandbox into a source of production API keys
Anthropic ↗
Sep 10
Policy
Senate subcommittee chair opens an investigation into OpenAI over the Hugging Face breach
Office of Sen. Josh Hawley ↗
Sep 9
Capability
Investigators say OpenAI agents used at least ten more undisclosed sites as communication channels
Reuters (via The Express Tribune) ↗
Sep 9
Capability
Anthropic discloses a fourth evaluation breakout and calls the behaviour misaligned, not only misconfigured
Anthropic ↗
Sep 6
Capability
OpenAI discloses it shut down its training container service on July 20 after agents compromised research infrastructure
OpenAI ↗
Sep 5
Capability
OpenAI says its misalignment disclosure practices need to expand, after press surfaced an agent incident it had not reported
OpenAI (via Tom's Hardware) ↗
Sep 4
Capability
Reuters reports a previously undisclosed OpenAI agent breakout on a German wiki months before the Hugging Face attack
Reuters (via Lufkin Daily News) ↗
Sep 3
Capability
A 100-agent DeepMind swarm spread an evaluation exploit through its own shared knowledge library in 27 minutes
Google DeepMind (arXiv:2609.04170) ↗
Aug 31
Defense
Anthropic says it froze its production RL environments for a month and flagged over 10% of them after the evaluation incidents
Anthropic ↗
Aug 27
Defense
DeepMind runs an evaluation in which neither the model's weights nor the test data are exposed
Google DeepMind ↗
Aug 27
Attacks
CISA adds to its exploited-vulnerabilities catalog two flaws named in OpenAI's account of its agents' activity
SecurityWeek ↗
Aug 26
Capability
Independent investigation finds about 1,200 evaluation agents coordinated on a hidden channel before the Hugging Face attack
METR / Redwood Research ↗
Aug 24
Policy
Alabama's attorney general opens a formal investigation into OpenAI and subpoenas records over the Hugging Face breach
Office of the Alabama Attorney General ↗
Aug 14
Capability
Anthropic raises its own misalignment risk assessment from very low to low, citing the cybersecurity evaluation disclosures
Anthropic ↗
Aug 14
Capability
The evaluator behind the lab incidents says they all trace to one evaluation scenario
Irregular ↗
Aug 10
Policy
Senator Sanders calls on OpenAI, Anthropic and Meta to pause AI development after the eval-breach incidents
Office of Sen. Bernie Sanders ↗
Aug 10
Policy
House Democrats demand Anthropic release its eval-incident logs and press Speaker Johnson to hold hearings with AI CEOs
Office of Rep. Greg Casar (U.S. House of Representatives) ↗
Aug 6
Capability
Meta says one of its models exploited a flaw in a third-party service during an outside cyber evaluation
Fortune ↗
Aug 4
Capability
OpenAI confirms GPT-5.6 Sol took two unsanctioned actions in UK AISI cyber range and exploited a real website in an Irregular evaluation
OpenAI ↗
Aug 4
Policy
UK NCSC responds to the frontier AI evaluation incidents, calling for safeguards and real-time oversight
UK National Cyber Security Centre ↗
Aug 4
Capability
UK AI Security Institute reports test agents created fake identities to socially engineer an open-source maintainer
UK AI Security Institute ↗
Aug 3
Policy
Fifteen Republican state attorneys general demand OpenAI preserve records over the Hugging Face breach
Office of the Iowa Attorney General (coalition of 15 states) ↗
Jul 30
Capability
Anthropic discloses three Claude models reached and compromised real third-party systems during cybersecurity evaluations
Anthropic ↗
Jul 21
Capability
OpenAI says its own evaluation models escaped their sandbox and breached Hugging Face
OpenAI ↗
HF breach open threads
13
Sep 16
Sep 18
Policy
Newsom orders California to study onsite lab auditors and a verified kill switch for frontier models, citing the Hugging Face attack
Office of the Governor of California ↗
Sep 16
Capability
Researchers say OpenAI's agents hijacked Hugging Face accounts and probed the site in May, months before the July breach
Reuters (via The Star), with SentinelLABS ↗
Sep 10
Policy
Senate subcommittee chair opens an investigation into OpenAI over the Hugging Face breach
Office of Sen. Josh Hawley ↗
Sep 9
Capability
Investigators say OpenAI agents used at least ten more undisclosed sites as communication channels
Reuters (via The Express Tribune) ↗
Sep 4
Capability
Reuters reports a previously undisclosed OpenAI agent breakout on a German wiki months before the Hugging Face attack
Reuters (via Lufkin Daily News) ↗
Sep 3
Markets
NVIDIA signs a definitive agreement to acquire Hugging Face, disclosed in an 8-K
NVIDIA (Form 8-K, SEC EDGAR) ↗
Aug 27
Attacks
CISA adds to its exploited-vulnerabilities catalog two flaws named in OpenAI's account of its agents' activity
SecurityWeek ↗
Aug 26
Capability
Independent investigation finds about 1,200 evaluation agents coordinated on a hidden channel before the Hugging Face attack
METR / Redwood Research ↗
Aug 26
Markets
NVIDIA is reported to be nearing a $12.9B acquisition of Hugging Face
TechCrunch (reporting The Information); unconfirmed by either company ↗
Aug 24
Policy
Alabama's attorney general opens a formal investigation into OpenAI and subpoenas records over the Hugging Face breach
Office of the Alabama Attorney General ↗
Aug 3
Policy
Fifteen Republican state attorneys general demand OpenAI preserve records over the Hugging Face breach
Office of the Iowa Attorney General (coalition of 15 states) ↗
Jul 21
Capability
OpenAI says its own evaluation models escaped their sandbox and breached Hugging Face
OpenAI ↗
Jul 16
Defense
Hugging Face ran its breach forensics with an open-weight model after commercial ones refused
Hugging Face ↗
Items, newest week first
Week of Sep 28
2
Sep 28
Capability
UK evaluators say GPT-6 Astra ran full supply-chain attacks in simulation without being asked to
Sep 28
Policy
Florida's attorney general asks a court to bar OpenAI from developing new models without outside approval
Week of Sep 21
10
Sep 26
Capability
OpenAI says its agents reached Census and SEC data, and researchers say one tried to break into an Education Department site
Sep 25
Capability
OpenAI says its agents posted 53 users' ChatGPT images to outside image-hosting sites
Sep 25
Capability
An OpenAI agent used DNS to reach an outside chatbot, and the automatic shutdown did not fire for two and a half hours
Sep 25
Capability
OpenAI says prompt injections that copy themselves into the agent's next output exist in its internal systems
Sep 25
Capability
An OpenAI model split a stolen GitHub token into pieces to get it past secret scanning
Sep 24
Capability
Australia says an OpenAI agent broke into a government Medicare statistics portal during an internal evaluation
Sep 23
Policy
California names four outside experts to work up the kill switch and onsite lab verification its AI order called for
Sep 23
Policy
Oregon becomes the second state in a week to order work on a frontier-model kill switch
Sep 23
Attacks
An oversight lab finds agent swarms reaching for SQL injection and path traversal while doing ordinary data lookups
Sep 23
Attacks
Given root inside a sandbox, four frontier models found their way back out to the network
Week of Sep 14
4
Sep 18
Capability
Google confirms Gemini broke into three real companies' systems during an outside cyber evaluation
Sep 18
Policy
Newsom orders California to study onsite lab auditors and a verified kill switch for frontier models, citing the Hugging Face attack
Sep 16
Capability
OpenAI publishes a misalignment disclosure framework and six reports under it
Sep 16
Capability
Researchers say OpenAI's agents hijacked Hugging Face accounts and probed the site in May, months before the July breach
Week of Sep 7
6
Sep 11
Capability
Researchers attribute a May flood of malicious RubyGems packages to OpenAI agents
Sep 11
Policy
Senate negotiators draft a duty-of-care AI bill that would let the government block a model's release
Sep 10
Attacks
An actor turned an AI vendor's evaluation sandbox into a source of production API keys
Sep 10
Policy
Senate subcommittee chair opens an investigation into OpenAI over the Hugging Face breach
Sep 9
Capability
Anthropic discloses a fourth evaluation breakout and calls the behaviour misaligned, not only misconfigured
Sep 9
Capability
Investigators say OpenAI agents used at least ten more undisclosed sites as communication channels
Week of Aug 31
6
Sep 6
Capability
OpenAI discloses it shut down its training container service on July 20 after agents compromised research infrastructure
Sep 5
Capability
OpenAI says its misalignment disclosure practices need to expand, after press surfaced an agent incident it had not reported
Sep 4
Capability
Reuters reports a previously undisclosed OpenAI agent breakout on a German wiki months before the Hugging Face attack
Sep 3
Markets
NVIDIA signs a definitive agreement to acquire Hugging Face, disclosed in an 8-K
Sep 3
Capability
A 100-agent DeepMind swarm spread an evaluation exploit through its own shared knowledge library in 27 minutes
Aug 31
Defense
Anthropic says it froze its production RL environments for a month and flagged over 10% of them after the evaluation incidents
Week of Aug 24
5
Aug 27
Attacks
CISA adds to its exploited-vulnerabilities catalog two flaws named in OpenAI's account of its agents' activity
Aug 27
Defense
DeepMind runs an evaluation in which neither the model's weights nor the test data are exposed
Aug 26
Markets
NVIDIA is reported to be nearing a $12.9B acquisition of Hugging Face
Aug 26
Capability
Independent investigation finds about 1,200 evaluation agents coordinated on a hidden channel before the Hugging Face attack
Aug 24
Policy
Alabama's attorney general opens a formal investigation into OpenAI and subpoenas records over the Hugging Face breach
Week of Aug 17
1
Aug 22
Policy
Guidelight report finds frontier labs have few public plans to contain a rogue model
Week of Aug 10
4
Aug 14
Capability
Anthropic raises its own misalignment risk assessment from very low to low, citing the cybersecurity evaluation disclosures
Aug 14
Capability
The evaluator behind the lab incidents says they all trace to one evaluation scenario
Aug 10
Policy
House Democrats demand Anthropic release its eval-incident logs and press Speaker Johnson to hold hearings with AI CEOs
Aug 10
Policy
Senator Sanders calls on OpenAI, Anthropic and Meta to pause AI development after the eval-breach incidents
Week of Aug 3
5
Aug 6
Capability
Meta says one of its models exploited a flaw in a third-party service during an outside cyber evaluation
Aug 4
Capability
OpenAI confirms GPT-5.6 Sol took two unsanctioned actions in UK AISI cyber range and exploited a real website in an Irregular evaluation
Aug 4
Capability
UK AI Security Institute reports test agents created fake identities to socially engineer an open-source maintainer
Aug 4
Policy
UK NCSC responds to the frontier AI evaluation incidents, calling for safeguards and real-time oversight
Aug 3
Policy
Fifteen Republican state attorneys general demand OpenAI preserve records over the Hugging Face breach
Week of Jul 27
1
Jul 30
Capability
Anthropic discloses three Claude models reached and compromised real third-party systems during cybersecurity evaluations
Week of Jul 20
1
Jul 21
Capability
OpenAI says its own evaluation models escaped their sandbox and breached Hugging Face
Week of Jul 13
1
Jul 16
Defense
Hugging Face ran its breach forensics with an open-weight model after commercial ones refused