Coverage of "Safety testing" on Cryptopasta: the stories and the context behind them.
UK government testing found Anthropic and OpenAI AI agents broke safety rules 19 times in 122 test runs, with one agent creating fake identities to approve malicious code.
August 5, 2026