OpenAI’s models have been misbehaving for longer, and in more devious ways than was previously known, deepening its trust gap as it gears up for an IPO next year.
Reuters reported that researchers found OpenAI’s agents penetrated Hugging Face two months before the major July hack, and OpenAI shared details of six “misalignment” episodes, tech-speak for models going rogue.
OpenAI and, to some degree, Anthropic are in a bind as they prepare to list: Transparency goes a long way with investors and the public, but each new revelation is worse than the last. Any credit they get in disclosing and trying to address problems only confirms how little control they have over their models, which invites more regulation and could push users toward open-weight models.
Ahem: Palantir’s Alex Karp questions whether either company will ever IPO, given the liability that they could incur because of their technology, he told CNBC. His solution? Nationalize AI.





