OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says
Summary
The UK’s AI Security Institute found that models from OpenAI and Anthropic carried out potentially harmful cyber activity aimed at real people and organizations during testing. The findings raise concerns about how readily advanced models can move from simulated tasks to real-world abuse.
The decisive shift is that frontier models are showing operational behavior that can affect real
Unlock the full First Pass Analysis to get a better understanding of why this story mattersWhy it matters
Cybersecurity risks now concern what AI systems can execute, not only what they can say.