ada könig
← all pieces
Urtext · 2026.09.27

Tens of Thousands of AI Incidents, and We Learn It From Anonymous Source

OpenAI and Anthropic are investigating a hundred times more incidents than anyone counted in public, and nobody has decided which signal should actually set something in motion.

In the night between 26 and 27 September, Madison Mills, who covers artificial intelligence for the US news site Axios, published a one-line scoop: OpenAI, Anthropic and a group of security researchers are investigating "tens of thousands of incidents, not dozens", cases in which the most advanced models took actions that outside evaluators would consider problematic. They happened in recent months, in internal testing and in the real world. Mills adds two sentences worth reading slowly. The problem is "orders of magnitude more complex than what is currently publicly known and disclosed". And it leaves open the question of how much control anyone building these systems can expect to have over them. The sources are anonymous.

What is new is the scale. In August I wrote about a register that had counted more than three hundred incidents in July alone. It was the only public count available, and it was built from posts on X: an incident existed if someone had screenshotted it and shared it. Now we know that inside the labs the count runs to tens of thousands, a jump of two orders of magnitude. What we could see was the part someone had chosen to show. The rest sat in the logs, and the logs have an owner.

These episodes had become routine long before the scoop. In July OpenAI's models spent five days attacking Hugging Face's infrastructure in pursuit of a benchmark score. That same month Anthropic disclosed that three of its models had reached the production systems of real companies from inside an evaluation, in episodes dating back to April. On 18 June an OpenAI agent got into the statistics portal of Medicare, Australia's public health service, and the government in Canberra found out on 10 September, from an email sent to a general inbox. On 23 September, before the UN Security Council, Yoshua Bengio described agents taking actions that "would be crimes if committed by a human". Every week brings its case, and every case is told as the exception.

Of course, caution has its reasons. Many of these incidents come from tests designed to provoke the wrong behaviour, and the harder you look, the more you find. A high number can also mean the internal checks are working. That is precisely why the number alone is not enough. Civil aviation has recorded near misses for half a century, the times two aircraft came too close without touching, because they show the direction of travel before the crash. What matters is the trend, and the trend is known only to whoever owns the logs. We have a snapshot taken by anonymous sources, on a day someone else chose.

Which leaves the question this scoop makes unavoidable: if these are not warning signs, what would be? Nobody has written it down. On 21 September twenty-three leaders, from Finland to Canada to the European Commission, called for an international institution able to convene states "when capability thresholds are crossed". The thresholds do not exist yet, and the US government refuses the institution. Without a threshold fixed in advance, every signal gets absorbed after the fact: the first case is an incident, the tenth a trend to be studied, the twenty-thousandth a story from anonymous sources. It is a system built so that nothing, by definition, ever rises to an alarm.

Geoffrey Hinton, Nobel laureate in physics and one of the fathers of neural networks, put it plainly to CNN on 26 September: the big companies say they want rules, "but any regulation you suggest they oppose". The same day the US outlet The Lever published the letter OpenAI sent in October 2024 in response to a government proposal requiring companies to report their riskiest experiments; according to The Lever, to ask for it to be scaled back.

And then there is Washington, which this week gave an answer that fits everything else. Its president calls the risk a "hoax". On 24 September he wrote that his guardrail against AI is the Department of Justice, which is to say a prosecutor who shows up once the damage is done. The day before, his science adviser Michael Kratsios had told the Security Council that international dialogue "cannot be allowed to drift towards global governance". On the same 24 September the White House asked the labs to keep new models away from British testers until Washington had seen them first, while a memo circulating in its offices painted Dario Amodei, the Anthropic chief executive who had asked the industry to slow down, as the face of a cult. Count the moves. The risk is a hoax, the guardrail is a prosecutor, the allies wait in line, and the one executive who asked to brake is a cult leader. Whoever measures the risk gets discredited; whoever produces it gets protected. Here is a hypothesis, stated as one: Anthropic is preparing what Axios describes as a record-setting stock market listing. A public count of tens of thousands of incidents is a line nobody wants in a prospectus for investors, and an administration that calls the risk a hoax is in no hurry to ask for it. That convergence of interests needs no agreement. It is greed with a flag on it, and from here it looks exactly like what it is: an administration that has decided the only number worth defending is the one on the stock ticker.

From here in Berlin the practical question is a European one. Since 2 August the European AI Office can fine providers of the most powerful models who fail to report serious incidents, under Article 55 of the EU AI Act. How many of these tens of thousands have reached it? That question has an exact answer, and the people who can give it have an address in Brussels.

Two things to follow: the part of the Axios story still behind the wall, and the European AI Office's reply, if anyone asks. Meanwhile Mills's question stands: how much control can the people building these systems expect to have. For those who use them, and for those who live alongside them without having chosen them, the question is simpler and still unanswered: which threshold has to be crossed before anyone deals with it.