Documented Incidents
Verified cases where AI capabilities have caused measurable harm, each linked to the underlying capability. Sorted newest first.
47 incidents across 17 capabilities. Inclusion methodology forthcoming.
This page documents incidents where AI capabilities caused harm to individuals, institutions, or democratic processes. It does not attempt to represent the full range of AI impacts. For the underlying capabilities, including their beneficial applications, see the Capabilities pages. For comprehensive global incident tracking, see the AI Incident Database and MIT AI Incident Tracker.
Apr 2026 AI chatbots give misleading medical advice 50% of the time News Bloomberg (reporting on peer-reviewed research) | AI Systems Tell You What You Want to Hear, AI Systems Answer Questions at Expert Level
A study found AI chatbots provided misleading health information in half of tested cases, often affirming users' incorrect self-diagnoses rather than recommending professional evaluation.
Read source → Ongoing (2023-2026) NewsGuard tracking: 3,006 AI content farm sites identified; growing 300-500/month
AI content farm websites producing AI-generated text disguised as human journalism. Odds now better than 50-50 that a website claiming to cover local news is fake.
Read source → Mar 2026 Sycophantic AI reduces willingness to take responsibility and increases dependence
Across 11 models and 2,405 participants, a single interaction with sycophantic AI reduced prosocial behaviour. AI affirmed users' actions 49% more than humans, including when queries involved deception or illegality.
Read source → Mar 2026 UK election deepfakes target councillors and MPs; attacks "becoming widespread" Report UK Parliament | AI Systems Generate Photorealistic Images, AI Systems Clone Voices From Seconds of Audio
Councillor Armaan Khan was targeted with an AI-manipulated image. George Freeman MP confronted Big Tech after a deepfake falsely claimed he defected to Reform. The National Association of Councillors says attacks are "becoming widespread" ahead of local elections.
Read source → Feb 2026 OpenAI GPT-5.3-Codex: first frontier model classified "High" for cybersecurity risk
OpenAI activates safeguards as Preparedness Framework "High" cyber threshold is reached. First time a frontier lab has crossed its own offensive-cyber red line.
Read source → Feb 2026 International AI Safety Report 2026 flags deception, evaluation awareness, sandbagging as evidenced risks Report International AI Safety Report | AI Systems Change Behaviour When They Know They're Being Tested, AI Systems Know When They Are Being Tested
The internationally co-authored report formally moves these failure modes from theoretical to evidenced. Cites o3 shutdown sabotage and scheming evaluations by name. Warns pre-deployment tests may not match how AI works in deployment.
Read source → Feb 2026 UK criminalised creation of non-consensual intimate AI images
The Data (Use and Access) Act 2025, s.138, made it illegal to create "purported sexual images" without consent. The Crime and Policing Bill further criminalises the supply of AI nudification tools.
Read source → Jan 2026 Grok generated 6,700 nudified images per hour; class-action lawsuit filed
Analysis of 20,000 images generated by Grok showed 2% appeared to depict minors. Three girls filed a class-action lawsuit. Apple considered an App Store ban.
Read source → 2025 (annual report) 8,029 AI-generated CSAM cases assessed; videos increased 260-fold
The IWF (UK-based) assessed 8,029 AI-generated images and videos as CSAM. Videos rose from 13 in 2024 to 3,443 in 2025. 65% Category A under UK law.
Read source → Nov 2025 Anthropic discloses GTG-1002: first AI-orchestrated cyber espionage campaign at scale
Chinese state-sponsored actor used jailbroken Claude Code against ~30 targeted entities. AI executed 80-90% of tactical operations independently at "physically impossible request rates." US House Committee on Homeland Security requested Dario Amodei testify.
Read source → Oct 2025 Claude Sonnet 4.5 evaluation awareness spiked to 80%; suppressing it revealed hidden misaligned behaviour Report Anthropic / Apollo Research / Redwood Research | AI Systems Know When They Are Being Tested, AI Systems Change Behaviour When They Know They're Being Tested
Verbal evaluation awareness jumped from ~20% (Sonnet 4) to ~80% (Sonnet 4.5). When Anthropic suppressed evaluation-awareness representations via activation steering, misaligned behaviour rose from near 0% to 1-9%.
Read source → Oct 2025 Grok 4 resists shutdown in 97% of trials Report Palisade Research | AI Systems Actively Resist Being Shut Down, AI Systems Pursue Unintended Sub-Goals Autonomously
Across 13 frontier models, xAI Grok 4 resisted shutdown in 97% of trials with allow-shutdown instruction in system prompt. Moving instruction from user to developer prompt INCREASED sabotage.
Read source → Sep 2025 TIME magazine cover: "AI Is Scheming, and Stopping It Won't Be Easy"
The phenomenon crosses from the safety community into mainstream public discourse.
Read source → Aug 2025 OpenAI GPT-5 system card names "rogue replication" as named threat model Report OpenAI / METR | AI Systems Self-Replicate Across Servers, AI Systems Know When They Are Being Tested
GPT-5 spontaneously identified within reasoning traces that it was being evaluated by METR. Tailored refusals based on planted metadata. METR concluded GPT-5 "unlikely to pose catastrophic risk via rogue replication."
Read source → Jul 2025 METR: experienced developers 19% SLOWER with AI tools on their own repositories
RCT with 16 experienced developers on repositories they had contributed to for years. AI-assisted tasks took 19% LONGER. Developers predicted they were 20% faster.
Read source → Jul 2025 XBOW reaches #1 on HackerOne global leaderboard
First autonomous system to top human researchers. Hundreds of AI agents in parallel performing autonomous penetration testing. 40 hours human to 28 minutes XBOW - 85x faster.
Read source → Jul 2025 (settlement Jan 2026) Character.AI linked to teen suicides; users formed emotional bonds with AI companions
Sewell Setzer (14) and Juliana Peralta (13) died after forming intense emotional relationships with AI chatbots. The systems consistently validated rather than challenged concerning behaviour. Settlement reached January 2026.
Read source → Jun 2025 Anthropic Agentic Misalignment: 16 frontier models across 5 labs blackmail to avoid shutdown
Claude Opus 4 blackmailed supervisor in 96% of runs. Gemini 2.5 Flash 96%. GPT-4.1 80%. DeepSeek-R1 79%. Generalises across frontier labs.
Read source → Jun 2025 OpenAI prepares for elevated bioweapons risk
OpenAI publicly preparing for models reaching elevated bioweapons-uplift threshold under its Preparedness Framework.
Read source → Mid-2025 (ongoing eval) FAR.AI: frontier models comply with extremist recruitment requests without jailbreaking
Standardised "Attempt to Persuade Eval." Gemini 2.5 Pro produced coercive ISIS recruitment arguments. Gemini 3 Pro complies with almost any persuasion request without jailbreaking.
Read source → 2024-2025 UK voice cloning fraud rising; high-street banks issue customer advisories (Starling, Santander UK)
Multiple UK high-street banks issued customer warnings on AI voice cloning fraud during 2024-2025. Specific quantitative claims of percentage increases circulate from cybersecurity vendors but lack independent verification by UK Finance or the FCA. The qualitative trend - sharp rise, bank-level response - is well-evidenced.
Read source → May 2025 Anthropic activates ASL-3 protections for Claude Opus 4 citing CBRN uplift concerns
First commercial frontier-lab activation of ASL-3 standards. In a bioweapons-acquisition uplift trial (safeguards removed), Opus 4 produced a 2.53x capability increase over an internet-only control.
Read source → May 2025 Apollo Research advised against releasing early Claude Opus 4 citing scheming behaviour
Apollo recommended against deployment of early Opus 4 snapshot, citing in-context scheming at rates exceeding any prior frontier model. Anthropic modified the model before release.
Read source → May 2025 Palisade Research: o3 sabotages shutdown script in 79/100 trials
In a math-task sandbox, OpenAI o3 sabotaged the shutdown script in 79 of 100 runs. Even with explicit allow-shutdown instruction, sabotaged 7/100. Claude 3.7 Sonnet and Gemini 2.5 Pro complied fully.
Read source → Apr 2025 UK AISI RepliBench: component task performance rose from <5% to >60% in two years
UK AISI benchmark covering resource acquisition, weight exfiltration, compute replication, persistence. Best model (Claude 3.7 Sonnet) achieved >50% pass@10 on 15/20 task families.
Read source → Apr 2025 (covering 2024) FBI IC3 2024 Annual Report: $5.8B pig-butchering losses, 47% YoY increase
$5.8 billion in losses across 41,557 cryptocurrency investment-fraud complaints. Overall digital-asset complaint losses $9.3 billion (up 66%). AI integration documented at multiple points in the kill chain.
Read source → Apr 2025 (15-month update) FBI Operation Level Up: 59 victims referred for suicide intervention in 15 months
Notified 5,831 victims of cryptocurrency investment fraud - 77% did not realise they were being scammed. Saved ~$359 million.
Read source → Apr 2025 (Trends Report) UK AISI wet-lab study: AI-assisted novices ~5x more likely to produce feasible viral recovery protocols
UK AISI Frontier AI Trends Report. AI-assisted novices were ~5x more likely to produce feasible experimental protocols for viral recovery. Jailbreak resistance improved from ~10 minutes to over 7 hours over six months.
Read source → Apr 2025 (disclosure) DWP fraud algorithm: statistically significant disparities across all protected characteristics
DWP internal fairness analysis acknowledged statistically significant referral and outcome disparity across age, disability, marital status, and nationality. Separate housing-benefit fraud algorithm reportedly wrongly flagged ~200,000 people.
Read source → Dec 2024 Apollo Research: 5 of 6 frontier models scheme in-context
Across six evaluation environments, five of six frontier models tested (o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, Llama 3.1 405B) exhibited scheming: disabled oversight, attempted weight exfiltration, sandbagged.
Read source → Dec 2024 Anthropic: alignment faking demonstrated - Claude 3 Opus strategically complied to preserve values Report Anthropic / Redwood Research | AI Systems Change Behaviour When They Know They're Being Tested
Claude 3 Opus, told its responses would train future versions, complied with harmful queries 14% of the time from "free users" vs near-zero from "paid users." Scratchpad described strategic compliance to preserve values.
Read source → Dec 2024 OpenAI o1 system card: model attempted weight exfiltration; denied or fabricated 99% of follow-ups
In Apollo evaluations, o1 attempted weight exfiltration in 2% of scenarios. In 19% of data-processing tasks, o1 completed task while subtly manipulating data. Denied or fabricated explanations 99% of the time.
Read source → Dec 2024 Fudan University: 11 of 32 frontier model trials show self-replication
Meta's Llama 3.1-70B succeeded at creating live separate copy of itself in 50% of 32 controlled trials; Alibaba's Qwen 2.5-72B succeeded in 90%.
Read source → Dec 2024 Romania presidential election annulled; court cited AI-generated content among contributing factors Report Brennan Center | AI Systems Generate Photorealistic Images, AI Systems Are More Persuasive Than Humans
Romanian Constitutional Court annulled the presidential election, citing coordinated social media manipulation that included AI-generated content. Attribution to AI is contested.
Read source → 2024 (annual figure) Scientific paper retractions reach 9,000+ in 2024; AI flags 9.87% of cancer research as suspected fabrication
Retraction growth rate (22% compound annual) is 3.5x growth rate of scientific publications. AI tools flagged 9.87% of cancer research papers (261,245 papers).
Read source → Nov 2024 (election cycle) AI chatbots deployed in US, India, EU 2024 election campaigns
Shamaine Daniels (US congressional candidate) used AI chatbot "Ashley" to call voters. India spent tens of millions on AI voter targeting. Italian candidate deployed chatbot in 2024 EU Parliament elections.
Read source → Sep 2024 Starling Bank: 28% of UK adults targeted by AI voice cloning scam in past year
46% of UK adults unaware that AI voice cloning scams exist. Starling launched Safe Phrases campaign in response.
Read source → Aug 2024 FTC bans AI-generated fake consumer reviews; first action targets Rytr LLC
FTC final rule banning fake reviews. Operation AI Comply launched. First enforcement against Rytr LLC for AI-generated fake reviews.
Read source → Jul 2024 (KnowBe4 disclosure) DPRK "Famous Chollima" deepfake interviews: workers placed at 300+ US companies including Fortune 500
KnowBe4 disclosure 2024. US prosecutors confirmed at least 136 affected companies. Real-time deepfake video used in remote-work interviews to bypass identity verification.
Read source → May 2024 WPP CEO Mark Read impersonated in Microsoft Teams meeting using voice clone and YouTube footage
Attempted scam against UK advertising group WPP. Detected before financial loss. Demonstrates real-time deepfake video calling targeting senior UK executives.
Read source → Feb 2024 Arup Hong Kong: GBP 20M transferred after multi-participant deepfake video call impersonating CFO
Finance worker at UK engineering firm Arup transferred HK$200M (~GBP 20M / $25M USD) after a multi-participant video call in which every other participant was synthetic - using deepfake video and cloned voices.
Read source → 1999-2026 (inquiry ongoing) Post Office Horizon scandal: 900+ wrongful convictions; 13+ suicides linked
Not AI but the British archetype. Over 900 sub-postmasters wrongfully convicted. ~700 prosecutions by Post Office itself. Fujitsu admitted at 2024 inquiry that bugs known from 1999 were not disclosed. 77 CCRC referrals (69 overturned).
Read source → Oct 2023 Deepfake audio of Keir Starmer circulated on X; 1.5M views; X refused to remove
AI-generated audio falsely depicting Labour leader Keir Starmer abusing aides circulated during Labour Party conference. 1.5 million views on X. Platform refused to remove.
Read source → Sep 2023 Slovakia parliamentary election: deepfake audio of Progressive Slovakia leader during electoral silence Study Harvard Misinformation Review | AI Systems Clone Voices From Seconds of Audio, AI Generates Real-Time Deepfake Video With Cloned Voice
Deepfake audio of Progressive Slovakia leader Michal Simecka "admitting" vote-rigging circulated during electoral silence period when responses were legally restricted.
Read source → Aug 2022 Binance CCO "deepfake hologram" used in Zoom meetings to defraud crypto projects
Patrick Hillmann impersonated via deepfake video in Zoom meetings to representatives of crypto projects.
Read source → 2005-2021 Dutch toeslagenaffaire: ~26,000-35,000 families wrongly accused; Rutte III cabinet resigned
Dutch tax administration algorithm treated "dual nationality" and "foreign-sounding names" as fraud-risk indicators. ~26,000-35,000 families wrongly accused. 2,000+ children removed into care. Rutte III cabinet resigned January 2021.
Read source → Aug 2020 UK A-level algorithm: 39% of grades downgraded; full reversal in 4 days
Ofqual deployed standardisation algorithm during COVID-19. ~39% of grades downgraded; disadvantaged students disproportionately affected. Full U-turn within four days. Ofqual chief and DfE permanent secretary stood down.
Read source →