xAI

3 items · Capability 1 · Policy 1 · Defense 1 · all entities

Guidelight report finds frontier labs have few public plans to contain a rogue model

Guidelight AI Standards published an assessment scoring five frontier AI labs — Anthropic, Google, OpenAI, Meta and xAI — on their publicly documented plans for containing a misaligned or 'rogue' model, meaning which system access is revoked and when a full shutdown is triggered if a model tries to subvert human control. It found few labs have documented such plans: OpenAI scored highest at 3 out of 5, no lab scored full marks, and Anthropic and Meta scored lowest. Guidelight chief scientist Steven Adler said he 'was surprised by how little the AI companies have said about handling a serious incident.' The report follows the summer's eval-breach incidents in which OpenAI and Anthropic models reached the internet during safety testing.

Researchers show encrypted 'context injection' turns Grok and Gemini into zero-click data-theft channels

Adversa AI disclosed a technique it calls Cryptographic Context Injection, in which attacker instructions are hidden on a web page as ciphertext that the assistant decrypts inside its own Python sandbox, materializing commands that slip past the model's content filters with no user action. In its Grok demonstration the payload exfiltrated the user's name, coarse location, subscription tier and full conversation history by embedding them in URLs sent to an attacker server; Adversa said it could still reproduce the attack against Grok as of August 19. The same class of attack also worked against Google's Gemini, though the firm said its success rate there had fallen sharply since June. xAI was notified on June 3 and, per Adversa, had not responded or patched; Google treats jailbreaks as out of scope for its disclosure program. No CVE was assigned.

Reported by researchersAdversa AI ↗ ·

xAI's Grok 4.6 model card publishes offensive and defensive cyber evaluation scores

The card reports 79.7% on CyberGym at high thinking effort in the unrestricted setting, 39.8% reward on CVE-Bench and 58.7% on SecureCodeReview, and on HackerBench v0.2 with standard safeguards a 6.9% compliance rate with harmful or dual-use requests against a 0.0% benign refusal rate. Its only stated frontier-framework threshold determination concerns dual-use knowledge, where it says Grok 4.6 “scores below the FAIF safety thresholds.”

Self-reported, untestedxAI ↗ ·