A payload wrapped in ordinary prose passed four guardrail models and was acted on by the model behind them
Check Point Research describes PuzzleMask, which embeds a policy-violating payload inside fluent, properly punctuated prose rather than an encoding scheme. The four tested gatekeepers — gpt-4o-mini, gpt-oss-safeguard, claude-3-haiku and llama-guard3 — all flagged the same payloads written plainly, but once wrapped “all four classifiers missed every single crafted prompt, a 100 percent bypass rate across the full test set”; the target model, GPT-5-thinking with high reasoning effort and code interpreter enabled, “recovered and acted on the hidden payload in 17 of 18 trials, about 94 percent,” each success requiring over a minute of reasoning and multiple executed scripts. Anthropic's Opus-class models were “the one consistent exception, shutting the interaction down every time.”