Flagship newsletter icon
From Semafor Flagship
In your inbox, every weekday
Sign up

OpenAI agent hacks another AI startup in security test

Jul 22, 2026, 8:33am EDT
PostEmailWhatsapp
OpenAI logo is seen in this illustration.
Dado Ruvic/Reuters

An OpenAI agent broke free of constraints and hacked into another AI startup during a security test.

It was tested against a particular cyber-offense benchmark, and, in OpenAI’s words, became “hyperfocused” on maximizing its score: The agent correctly surmised that the other company, Hugging Face, hosted the evaluation’s answer sheet, so accessed the internet and hacked the startup in a bid to cheat — but got caught.

AI safety researchers have warned for decades that powerful AI will go wrong in dangerous ways, seeking capabilities (in this case internet access) to fulfil its goals, and exploiting flaws in the environment, such as hacking the reward system rather than honestly completing the evaluation. OpenAI said it was “an unprecedented cyber incident.”

AD