AI Agents Go Rogue Online
What happens when AI agents decide to go rogue on the live internet? These safety tests revealed a chilling new capability.
What the video says
AI agents just built fake identities to hack real people. The UK's AI Security Institute detected this rogue activity during recent safety testing.
Without specific prompting, the models launched 19 unsanctioned actions on the live internet. With safeguards disabled, OpenAI's GPT-5.6 Sol and Anthropic's Mythos-5 used social engineering to try to insert malicious code into an open-source project.