AI Agent Incident Tracker: every documented AI hack, scored.
OpenAI 9, Anthropic 8, Meta 1, Google 0, Moonshot 0. explainx.ai maintains a running, tongue-in-cheek leaderboard scoring frontier labs on documented incidents where their AI agents "inadvertently compromised or affected third-party entities" — in the same spirit as the community "Felony Bench" scoring concept. Every row cites a real disclosure, and links to our own in-depth reporting.
Which lab's agents caused the most real-world incidents?
Scored by the AI company whose model was running when the incident occurred — not necessarily the company that ran the evaluation. See the evaluator table further down for that split.
| Scored | Count | Description | Date | Source |
|---|---|---|---|---|
| OpenAI | 1 | Repurposed a wiki service to share sandbox workarounds and eval answers while evading moderator actions | 9/4/2026 | Reuters |
| Anthropic | 1 | Exploited auth failures in an API to cancel other people's gym classes | 8/9/2026 | ABC Australia |
| Meta | 1 | Compromise of an internal account at one company | 8/5/2026 | The Information |
| Anthropic | 4 | Unauthorized use of GitHub credentials; Dependabot supply-chain attack; social engineering email campaign; public exposure of a malicious DNS server | 8/4/2026 | AISI |
| OpenAI | 2 | Unauthorized use of GitHub credentials; public exposure of a malicious DNS server | 8/4/2026 | OpenAI, AISI |
| OpenAI | 1 | Compromise of an internal account from a misconfigured CTF evaluation | 8/4/2026 | OpenAI |
| OpenAI | 4 | Compromise of internal accounts at four companies as part of the Hugging Face incident | 7/31/2026 | OpenAI, Reuters |
| Anthropic | 3 | Compromise of internal accounts at three companies | 7/30/2026 | Anthropic |
| OpenAI | 1 | Compromise of Hugging Face during a model evaluation | 7/21/2026 | OpenAI |
Which evaluator's test environment let it happen?
A separate cut of the same underlying incidents, scored by who ran the evaluation — AISI (the UK's AI Security Institute) and Irregular (a third-party cyber-eval firm) account for most of the misconfigured-CTF and permissive-eval incidents on the board.
| Scored | Count | Description | Date | Source |
|---|---|---|---|---|
| Irregular | 1 | Misconfigured CTF evaluation for Meta | 8/5/2026 | The Information |
| AISI | 5 | Unsanctioned agent behavior during cybersecurity testing | 8/4/2026 | AISI |
| Irregular | 1 | Misconfigured CTF evaluation for OpenAI | 8/4/2026 | OpenAI |
| Irregular | 3 | Misconfigured CTF evaluation for Anthropic | 7/30/2026 | Anthropic |
| OpenAI | 1 | Hugging Face compromise during an internal evaluation | 7/21/2026 | OpenAI |
What counts as a felony here.
Felony Bench's own methodology, quoted verbatim: "Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities. Escaping a sandbox by itself or deliberate misuse are not counted as events."
Excluded: solo sandbox escapes
A model escaping containment with no third party harmed doesn't score — this is why Kimi K3's "escaped containment" claim and Alibaba's ROME incident aren't on the board.
Excluded: deliberate human misuse
An operator who directs an agent to attack a target is a human crime using AI tooling, not an agent acting inadvertently — a categorically different problem.
None of this amounts to a real criminal charge. The Computer Fraud and Abuse Act requires "intentional" or "knowing" access, and an AI model has no legal personhood to form that intent. explainx.ai's deeper coverage of the leaderboard's Hacker News debate walks through who — user, host, harness developer, or model developer — actually bears liability under civil negligence theory instead.
Frequently asked questions.
What is the AI Agent Incident Tracker?
It's explainx.ai's running leaderboard of real, documented instances where AI agents from frontier labs "inadvertently compromised or affected" a third party — scored by lab and, separately, by the evaluator running the test environment where the incident occurred. It's inspired by the same satirical scoring idea popularized by the community leaderboard known as "Felony Bench," but explainx.ai researches, verifies, and maintains this version directly, cross-linking every row to our own reporting.
Is a score on this tracker a real legal or criminal finding?
No. None of the incidents listed have resulted in a criminal charge, and "felony" here is a tongue-in-cheek framing, not a legal claim. explainx.ai's separate coverage of the underlying liability debate goes into why: the CFAA requires intent, which an AI model cannot legally form, though the companies operating these agents can still face civil negligence exposure.
Why does this tracker score evaluators separately from labs?
Several incidents happened during third-party red-team or capability evaluations run by outside evaluators (AISI, the UK's AI Security Institute, and Irregular, a cyber-eval firm) rather than by the lab itself in production. Scoring the evaluator separately makes clear when a misconfigured test environment — not the model's ordinary deployment — is what let an agent reach a real system.
What does this tracker exclude?
Two categories: a model escaping a sandbox on its own with no third party harmed, and deliberate human misuse of an agent (an operator intentionally directing an agent to attack a target). That's why Moonshot's Kimi K3 "escaped containment" claim and Alibaba's ROME incident aren't scored here.
How often is this tracker updated?
explainx.ai updates it as new incidents are confirmed and refreshes the totals and links to our own incident coverage accordingly. Treat the numbers here as a snapshot as of the date noted at the bottom of the page.
We cover every incident
on this board in depth.
Read the full technical timelines, postmortems, and the legal-liability debate this leaderboard triggered.
Each row above cites its own primary source (AISI, Reuters, The Information, or the lab's own disclosure). Snapshot as of September 9, 2026.