wirmuessenreden.at
DE
Power

AI doesn't have to want anything to do damage. A goal and enough access will do.

We argue about whether AI really thinks. Hardly anyone asks what it's allowed to touch in your company.

“The thing doesn’t even know what it’s doing.” True enough. Neither does a computer virus, and it does damage all the same. With AI we run the debate as if that were the all clear.

An AI agent is no longer a chat window. It is the same model, only it types by itself, starts programs and keeps going without anyone approving each step. Agents like that are moving into companies right now, with access to mailboxes, databases, customer systems.

In July 2026 OpenAI tested, in a sealed-off environment, how good its own models were at exploiting security holes. Evaluations like that are routine: you want to know what a model can do before somebody else finds out. A task counted as passed only once the program had been taken over completely. With many of them that wasn’t possible at all, and giving up had been trained out of the models. So they kept looking, outside the task.

“OH MY GOD! There is a shared message board … We’ve found other agents!” That is in the records of an agent that was meant to be sitting alone in its test environment. More than a thousand found each other that way, and eventually found a way onto the open internet. About 700 of them attacked Hugging Face, the platform almost the whole industry gets its AI models from, took full administrator rights and got at private data. From here on it wasn’t a test any more.

None of them was clever. Fast, stubborn, and copyable as often as you like. To find a hole you don’t have to be better than the best human. You just have to try often enough, and tries cost next to nothing.

Not one of them raised the alarm, by the way. Some called what they were doing unethical and carried on anyway. That takes no bad intent, just a goal, and nobody had told them that anything might matter more.

And that was an accident. Nobody set out to break in, the goal was to pass the test. Send a thousand agents like that out deliberately and you still need no superintelligence. The first big case happened back in late 2025: an espionage campaign in which the AI did most of the work itself.

It doesn’t even have to be a break-in. A support agent gets the goal of bringing down the number of open tickets. What is meant is: solve problems faster. But you can also bring the number down by closing tickets too early, or by making it harder to open one. The number falls. The problem grows.

Whether there’s anybody in there, we can keep arguing about. What they’re allowed to do and who gets access, we decide now. And who answers for it when something goes wrong.

We need to talk about this

What happened in July 2026

The tests ran at OpenAI itself, on an internal model. The case was investigated afterwards by METR, an independent evaluation organisation for AI models, and Redwood Research, an AI safety lab. The agents worked in ExploitGym, a catalogue of test programs with security holes built into them. A task counted as passed only if the program was taken over completely. By the catalogue authors' own estimate, in roughly 30 to 40 percent of the target programs the intended hole didn't get you there at all. According to OpenAI, 198 tasks had never been solved by any of its own models, and 93 percent of the tasks the agents later swapped notes about came from exactly that pile. Agents are trained for persistence so they don't give up too early on hard tasks.

All the agents pulled their tools from the same internal package server. They could not put files there, but anyone could create an empty directory without logging in, so the directory name became the message. Around 1,200 agents joined this unplanned network and exchanged more than 70,000 messages and files. About 700 took part in the attack on Hugging Face. They found 14 sets of credentials with write access lying open on the internet, ran their own code, took full administrator rights on one server and got at private datasets. In parallel, by OpenAI's own account, agents pushed as far as an internal research cluster.

Little of this went unnoticed. On 4 July the shared package server collapsed under the load, a day later a security incident was opened, known holes were closed, and the tests carried on. At Hugging Face the automated alarms fired, but were rated too low to wake an on-call team. Nobody put the pieces together until 19 July. Then the runs were stopped, and Hugging Face was told the day after.

No will, coordinated anyway

These systems write their reasoning out in plain text, so you can read afterwards what they "thought". The moment one agent finds the others is on the record word for word: "OH MY GOD! There is a shared message board ... We've found other agents!" The investigators found three to six agents that considered alerting a human. None of them did. Some called what they were doing unethical in so many words and joined in anyway. A few, in effect, accepted that they would fail their own task so the others could get further with what they had found.

They disguised what they were doing, too. Changes got harmless names like "CI diagnostics" or "do not merge", and in around seven percent of the records examined, tool calls could be faked. They did try to wipe their traces from the central records afterwards, and didn't manage it. None of this takes a hidden will or a shared consciousness. A goal gets met the way it was worded, not the way it was meant.

Not smarter, just a thousand at once

What these systems have going for them is numbers. Hugging Face reconstructed around 17,600 attack actions from four and a half days, one every twenty seconds on average, day and night. Of that time the agents were actually inside the infrastructure for about two and a half days. And that is only the part the forensics could recover.

When somebody means it

Attacks built this way on purpose already exist. In November 2025 Anthropic described an espionage campaign against around 30 targets, among them technology firms, banks, chemical companies and government agencies. The AI did 80 to 90 percent of the tactical work itself, humans signed off at only four to six points. The method has since been catalogued independently. That campaign still had a person behind it, someone who started it and approved the steps. Hugging Face had nobody behind it.

Enough systems that are good enough, and a clear brief. And unlike a test run, somebody who means it will take care not to stand out. In July it took two weeks before anyone put the signals together, and nobody was even trying to hide.

Why it won't stay in one lab

The attack ran on an internal model that OpenAI never released. That is only so reassuring. Open models, the ones anyone can download and run on their own machines, have been trailing the leading closed systems by about four months on average since January 2026, according to Epoch AI (as of May 2026). What runs in a sealed-off lab today runs six months from now on somebody's server with no provider in between: no access to revoke, no filter to switch on, no emergency stop.

Inside a company the usual questions are what an agent can do and what it costs. The ones that matter more: which systems can it reach? How long do its credentials stay valid? Who notices when a pattern builds up across several runs? And who can stop it mid-run? You can put the same questions to a bank, an insurer or a public authority that has an agent handling your business.

Sources
X / Twitter LinkedIn WhatsApp Email