XBOW

2 items · Capability 2 · all entities

An autonomous agent found three critical Microsoft remote-code-execution flaws

XBOW reports its agent found CVE-2026-32194 and CVE-2026-32191, command injection in Bing image-processing pipelines, and CVE-2026-21536, an unrestricted file upload, each rated CVSS 9.8, reaching NT AUTHORITY\SYSTEM on production Bing image-processing workers running Windows Server 2022 and uid=0 on Linux workers across multiple hosts and network ranges. XBOW says the findings were made with no human in the loop, and that Microsoft's acknowledgements list it as the finder for all three.

Confirmed by orgXBOW (Microsoft credited the findings) ↗ ·

XBOW publishes cross-model offensive-security comparison placing GLM-5.2 and Muse Spark 1.1 near frontier models at lower cost

XBOW ran black-box testing against vulnerable open-source applications across Muse Spark 1.1, GLM-5.2, GPT-5.5, Mythos, Opus 4.6, GPT-5, Gemini models and Grok 4.5. It reported Mythos as strongest, GLM-5.2 falling between GPT-5 and Opus 4.6, and Muse Spark 1.1 landing just below Opus 4.6, concluding that 'good-enough offensive capability is getting much cheaper, and that changes the threat model.'

Self-reported, untestedXBOW ↗ ·