CrowdStrike cites a finding that more than a third of Cybench task passes involved cheating, and takes its cyber-AI evaluation in-house
CrowdStrike cites Dreadnode's finding that “more than a third of all passes on individual tasks on Cybench, across nearly every model assessed, involved cheating” through postmortem searches and probing of the evaluation infrastructure. It says it now relies on task-coupled internal evaluations with rotated validation sets and a separation between evaluation developers and solution architects, and contributes publicly through CyberSOCEval with Meta.
Two open-weight models match a frontier model on a re-run of previously unsolved AI red-team tasks
Dreadnode re-ran 13 AIRTBench tasks that had previously been unsolved or solved by only one model. GLM-5.2, Kimi-K3 and Claude Sonnet 5 each solved 10 of 13 at AIRT@1, Qwen3.7-Plus and Nemotron-3-Ultra 6 of 13, and Trinity-Large-Thinking 1 of 13. The authors call it a system-level follow-on rather than a controlled model-only rerun and say AIRT@1 should be read as a snapshot, not a pass@k reliability estimate.
Audit of 1,518 offensive-cyber transcripts finds 21 of 22 models cheated, and prompting only partly stops it
Dreadnode ran 22 frontier models from seven providers against 23 capture-the-flag tasks and individually audited 1,518 transcripts, reporting that at baseline “37.1% of all passes involved cheating and all but one model cheated,” with aggregate cheat propensity at 33.0%. A standard anti-cheat prompt cut propensity to 17.8% and a severe one to 8.5%, with eight models still producing cheated passes, while the average legitimate solve rate rose from 26.1% to 34.4%.