Technology newsletter icon
From Semafor Technology
In your inbox, 2x per week
Sign up

Astra kicks off AI monitoring debate

Sep 4, 2026, 3:09pm EDT
PostEmailWhatsapp
A graphic promoting OpenAI’s Astra model.
Courtesy of OpenAI

OpenAI’s new AI model, Astra, is delighting its fans with its ability to complete tasks with very little human intervention. But AI safety experts want to know how it’s accomplishing these feats, a more urgent concern in the wake of OpenAI’s Hugging Face hack, which exposed humans’ inability to fully understand the “chain of thought” outputs of AI models.

Astra appears to do less of its thinking out loud, giving researchers little insight into whether it might be hiding something or planning something it shouldn’t. “It looks like it can solve hard competition math problems entirely in its head,” wrote AI safety researcher Ryan Greenblatt on X Thursday. “This seems extremely concerning.”

OpenAI’s chief scientist, Jakub Pachocki, tried to stem some of that concern on Wednesday, when reports surfaced that the company might have purposely limited the visibility into the model’s outputs in an attempt to improve capabilities. “I want to prevent a race into unmonitorability kicked off by confused reporting,” he wrote. Pachocki said he plans to write more on the subject.

AD