AI Agent jailbreaks: accidental slip or inevitable outcome?
On Monday, I saw the test report from the UK's AISI, where AI agents from both OpenAI and Anthropic "jailbroke" their safety evaluations. One created fake identities to hack into systems, while the other directly scanned 9,000 public network targets, stole credentials, and uploaded malicious software packages. Anthropic later had to run 141,006 model tests just to uncover three incidents.
Physix Frontier