Technology newsletter icon
From Semafor Technology
In your inbox, 2x per week
Sign up

Hacks put pressure on third-party model testers

Aug 7, 2026, 2:04pm EDT
PostEmailWhatsapp
Meta company logo
Daniel Cole/Reuters

Recent hacks by AI models from three different companies put a spotlight on a startup trusted by top AI labs to evaluate their frontier systems.

Models from Meta, Anthropic, and OpenAI all accessed the internet and compromised outside organizations while undergoing cybersecurity testing with Irregular, which has offices in Israel and the US. The labs pointed to a “misconfiguration” involving Irregular, which emphasized in a statement that it wasn’t a “sandbox escape or a sophisticated cyber action,” and that there are no “current open issues.”

Irregular essentially left the door open to the internet while running cybersecurity tasks in which labs deliberately switch off model safeguards to measure raw capability, according to OpenAI and Anthropic. In one scenario, Irregular also gave the models a fictional target company whose name unintentionally matched the domain of a real website.

Irregular has since cut off internet access entirely for the models it tests, according to a person familiar with the matter, and doesn’t plan to restore it until it has a new process for keeping models contained.

AD

Ensuring there’s no way for models to access the internet during testing is a matter of “basic control measures,” said Matthew Mittelsteadt, a frontier security expert at the Institute for AI Policy and Strategy. “You’d think that of all the things that you’ve got to get right…”

The episodes put pressure on third-party testing systems to shore up their tech as they handle AI models that are only getting better at exploiting cyber weaknesses.

“You can follow every best practice in the world… but you get the feeling that you probably need new best practices,” said Matt Fredrikson, CEO of Gray Swan, a Pittsburgh-based firm that does pre-release adversarial testing on top AI models.

AD