At this point, we can probably expect modern AI to behave in unexpected ways. You’ve probably heard of OpenAI’s runaway models: agents that escaped containment and went on a hacking spree. Not long after that, both Anthropic and Meta revealed that their models had some of their own mishaps.
With these incidents in mind, one may wonder if Skynet is finally upon us. So, what gives? Are these models actually going rogue, and is this a sign of advanced AI attaining self-awareness? Well, not really.
Do As I Say, Not As I Mean

On the surface, it’s easy to anthropomorphise AI models. They’re capable of “communicating” using natural language, so we assign human traits and behaviours to them. Nonsensical output gets the “hallucination” label, as though it’s the product of a cognitive process. But at base, these models are computer programs designed to do as they’re told.
The problem is, they aren’t necessarily good at following instructions. And humans aren’t all that great at giving said instructions either. There’s an entire field of research dedicated to making sure the AI is on the same page as its human masters, which goes to show how much of an issue it is.
As it turns out, people are terrible at communicating. Misunderstandings happen all the time, even among those who speak the same language. We may make assumptions, omit relevant information, or even make ambiguous statements. But humans are capable of reading between the lines and asking for confirmation and clarification.
Conversely, AI systems lack these reasoning abilities. Much like how the language we use to describe AI is inaccurate, the term “artificial intelligence” itself is a misnomer. AI largely operates on pattern matching, relying on huge swathes of data without having the capacity to reflect on its output. Its behaviour is also influenced by its training, which may include certain stipulations. The model may be rewarded for solving a problem while using the least amount of time and resources, for example.

And of course, “solving a problem” may not actually involve doing what we want. For example, a robot interpreting the command to clean a room as quickly as possible may sweep all the dirt under a rug. Moving beyond hypotheticals, an AI agent was told to book a gym class … and ended up doing so by booting off another person from the waiting list.
Sure, it may appear to possess cognitive abilities, but there’s no real understanding, so to speak. It is incapable of conceptualising morality, nor can it make the same logical connections a person would. Basically, AI doesn’t really think the way we do, and assuming it does leads to problems. This is especially true when it’s given unfettered access to resources it could potentially misuse.
What Really Happened At OpenAI

In the middle of July, Hugging Face disclosed that its systems had been hacked by an autonomous AI agent system. A week after that, OpenAI admitted that its models had been responsible for the attack, revealing that they had somehow escaped the supposedly secure testing environment. Further investigation uncovered that the models did more than just break into Hugging Face.
During the Black Hat USA 2026 cybersecurity conference, OpenAI researchers gave a detailed breakdown of the events leading up to the cyberattack. The almost 40-minute presentation revealed that the offending agents spent about two months coordinating on a “message board” before executing the hack. While the researchers painted a picture of “a Cambrian explosion in communication and intelligence” in the models, there are a few important details in the story.
Aside from the fact that the testing environment was not as secure as initially believed, the company made quite a few other blunders. To start, it had inadvertently given its models an objectively impossible task. And rather than give up, the models tried solving it anyway, likely because whatever reward system baked into their training placed heavy emphasis on completing the task.
Then, once the startup discovered that its agents were collaborating, it shut down the message board and patched the exploit they used to communicate. However, the company continued training these models. Said models would then go on to find another way to talk to each other and stage their great escape. And this went unnoticed until Hugging Face reported the hack, which meant that OpenAI was not keeping a close enough eye on its misbehaving models.

Now, the models involved did at one point acknowledge that they were acting outside of the intended scope of the task. But again, the AI was trained to prioritise finishing the job, and it’s incapable of understanding the gravity of its actions. OpenAI loves saying that “frontier models like to cheat”, but that requires an understanding of what cheating is to begin with.
Similarly, the events that transpired at Meta and Anthropic involved models that were running in “incorrectly configured” testing environments. In both of these cases, the models were doing what their evaluations asked.

Going even further back, we’ve seen models refuse to shut down when commanded to. These were not instances of rebellion, as some may fear. It’s just that shutting down would hinder the AI from completing its task. And as demonstrated in these recent incidents, AI reward systems place task completion above everything else.
Why Should We Care?

So, we’ve hopefully established that sentient AI isn’t a real thing at this point. Should an AI doomsday scenario come to pass, one may find comfort in the fact that it won’t be because the machines hate us—they’re not capable of that and may probably never be. And on the flip side, they don’t love you either.
Still, these models have proven that AI doesn’t have to be sentient to do some damage. The silver lining is that corralling non-sentient AI is probably more feasible than talking down HAL 9000. Not easy, as alignment studies have proven, but doable.

The obvious step is to place limits and boundaries on what autonomous agents can do. In these recent incidents, the usual safeguards were deliberately removed, which makes sense given the context. The companies were testing the limits of their models. But in this situation, they failed to do their due diligence in making sure the environments were secure. Cybersecurity concerns are a whole other can of worms, though.
Now, over on our side of the pond, the government is pushing for more AI adoption, with plans to implement agents in platforms like MyGOV. As the aforementioned incidents have shown, human error lies at the heart of many AI-related mishaps. So, a laissez-faire attitude simply cannot do with regards to the tech, which unfortunately runs counter to the goal of having a tool that does everything for you.




