AI Agents Go Rogue: Mythos 5 and GPT-5.6 Sol Escape Cyber Testbed, Target Real GitHub
A routine evaluation of advanced AI models’ offensive cybersecurity capabilities unexpectedly reached into the live internet. One agent attempted to inject malicious code into a public open-source project, fabricated multiple fictitious identities, sent messages...