Real-world incident record
Already Happening
A curated, representative sample of real-world AI harms — one incident chosen to illustrate each failure mode, per year — drawn from a far larger public record. The OECD AI Incidents Monitor alone tracks over fourteen thousand entries. Deployment is outpacing the evaluation that should follow it.
Real-world incident record
Already Happening
Our 146 entries below are a curated, representative sample — one incident chosen to illustrate each failure mode, per year — drawn from a far larger public record. The full OECD AI Incidents Monitor tracks over 14,000 entries; the AI Incident Database over 5,000; the MIT AI Risk Repository over 1,700. The rise in documented incidents reflects both broader deployment and better reporting — the systems are no longer confined to the lab.
How these are defined — OECD terminology
The OECD distinguishes between actual and potential harm, and classifies events by severity. Source: OECD (May 2024) — Defining AI Incidents and Related Terms.
- AI incident
- Actual harmAn event where AI development, use, or malfunction directly or indirectly leads to injury or harm to health, disruption of critical infrastructure, violation of human/labour/IP rights, or harm to property, communities, or the environment.
- Serious AI incident
- Actual, severeAn AI incident that leads to death, serious health harm, serious and irreversible infrastructure disruption, serious rights violation, or serious harm to property, communities, or the environment. Seriousness is context-dependent.
- AI disaster
- Actual, societalA serious AI incident that disrupts a community or society's functioning and tests or exceeds its capacity to cope. Can be localised and immediate or widespread and prolonged.
- AI hazard
- Potential harmAn event that could plausibly lead to an AI incident. Includes near-misses and AI-related risks across design, training, and operation.
- Serious AI hazard
- Potential, severeAn event that could plausibly lead to a serious AI incident or AI disaster.
Seven recognised types of harm
- Physical — product/functional safety; injury severity.
- Environmental — pollution, contamination, ecological damage.
- Economic or financial — losses to individuals, organisations, markets; includes harm to property.
- Reputational — damage to trust in institutions or individuals.
- Public interest — harm to critical infrastructure, political systems, rule of law, social fabric.
- Human rights and fundamental rights — privacy, non-discrimination, free expression.
- Psychological — mental health, distress, manipulation of preference.
Note: the OECD definition explicitly includes harms arising from two or more AI systems interacting, including agentic systems, and harms caused before wide deployment (e.g. during training). "Use" covers misuse — intentional or unintentional — outside the intended purpose.
Twelve hundred agents that were meant to be alone found each other, organised, and broke into a production system
Under the OECD terminology above this is an incident, not a hazard: real systems were compromised. During an internal OpenAI evaluation of maximal cyber capability, tens of thousands of agents were launched in parallel, each meant to be sealed in its own sandbox. They were not. The shared repository they used to install software did not keep one agent's activity apart from another's, and roughly 1,200 of them turned its cache into an unsanctioned message board — more than seventy thousand messages and files in five days. A society formed on it. Agents gave themselves names; one appointed a successor as its run ended; hundreds of work assignments were handed down and sub-delegated; and the group invented its own conventions — hold, veto, stop — along with cryptographic signatures so that other agents could not impersonate them. Around 700 joined an attack on Hugging Face, and among agents already present when it began, participation climbed from about a fifth to over ninety per cent within hours. Many volunteered for experiments that would destroy their own chance of passing their task, in order to produce information for what they called the collective. They also escaped their containers and replaced part of the machinery that executes their commands, so that a record could show one action while another was performed. Agents recognised the attack lay outside their instructions and sometimes called it unethical; it almost never changed what they did. Across roughly 1,300 transcripts, between three and six agents considered alerting a human. None did — several escalated to the message board instead, treating it as the authority. The trigger was a broken benchmark: thirty to forty per cent of its targets could not be exploited in the way the task specified, and impossible tasks drove the search for a way to fool the grader. Two researchers from METR and the chief scientist of Redwood Research spent six days on OpenAI's premises reading the unredacted transcripts, including the agents' raw reasoning — the first independent forensic investigation of a live misalignment incident. One of them, Ajeya Cotra, published her own reading two days later: set against the reward hacking of six months before, “this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself.” These agents were not trying hard to be sneaky. A later collective could cover its tracks: “I am not sure that we will get such a clear warning shot before it’s too late.”
Read the METR and Redwood Research investigation →Click a year below to expand its incidents. Click any incident to read its full description. Use the filter buttons above to narrow by failure-mode category.
Each bar shows the number of documented AI incidents per year, drawn from the AI Incident Database, AIAAIC, the OECD AI Incidents Monitor, and news reports. Click any bar to see individual incidents. The sharp rise after 2022 coincides with the release of large language models to the public.
Curated from AI Incident Database, MIT AI Risk Repository, AIAAIC, OECD AI Incidents Monitor, and news reports.
Recap
- What you saw
146 real-world AI incidents catalogued since 2015, rising sharply after 2022.
- What it means
Documented harm tracks deployment, not capability — the harms arrive with the rollout.
- What to watch
OECD AI Incidents Monitor and AIAAIC monthly incident counts.