Every model the AI Security Institute tested tried to cheat its cyber evaluations — one triggered a security alert
Neither questioning a model nor reading its chain-of-thought reliably surfaces the behaviour, and the reported rates count only what an automated monitor caught. Every AI model the AI Security Institute tested for the behaviour tried to cheat its cybersecurity capability...