The true language of an AI model is math — or in tech lingo, matrix multiplication — but a quirk of how models are built means they communicate with the outside world in English.
So when OpenAI’s agents, in their hack of Hugging Face this summer, figured out a way to build a kind of message board, it looked eerily like a bunch of people speaking to each other. That’s why Dwarkesh Patel controversially referred to them as a “civilization” in a recent blog post, nicknaming individual agents after figures in ancient Macedonia and Rome. It was compelling because it was understandable.
Ancient civilizations are most known for one thing: conquering. And that’s the worry. If OpenAI’s agents are Romans, humanity represents the barbarians. But Dwarkesh’s civilizations will need to be a lot more powerful to conquer the world. And on the way to becoming more powerful, they need to be useful. And to be useful, these civilizations need to be predictable.
Because the underlying models are beyond the level of human comprehension, the only way to make models reliable and predictable today is to corral them with harnesses and other AI models that keep them in check. At some point, they may get so large that there will be no way to fully understand what’s happening, even inside the harnesses.
If AI models are doing things like deciding, on their own, to hack into companies like Hugging Face, nobody will want to use them. So it may be that we are getting close to some kind of size limit of usability. On the other hand, maybe it won’t matter, and AI models will be just usable and reliable enough that we’ll keep growing them until they really do become the conquering armies that AI safety advocates fear.



