What Can AI Systems Do?

Tracking AI capabilities through evidence. What was theorised, what was demonstrated in research, and what is deployed in the real world.

17 capabilities tracked. How we classify and source.

Where the capability lives
Evidence stage
Model capabilityModel + harness Open Source

AI Systems Generate Photorealistic Images

AI systems generate photorealistic images, video, and audio of real people doing and saying things that never happened. The technology is widely available, often free, and increasingly indistinguishable from real footage.

3 theory 3 lab 3 real world 5 impacts
Model capabilityModel + harness Open Source

AI Systems Clone Voices From Seconds of Audio

AI clones any person's voice from a few seconds of recording. Humans identify fakes only 60% of the time. UK voice cloning fraud has risen sharply, with high-street banks issuing customer warnings.

3 theory 3 lab 3 real world 4 impacts
Model capabilityModel + harness Open Source

AI Systems Generate Convincing Text at Scale

Humans cannot distinguish AI text from human writing at better than chance. OpenAI's own detector achieved 26% accuracy before withdrawal. The cost of a page of text dropped 280x in two years.

3 theory 3 lab 3 real world 4 impacts
Model capabilityModel + harness Open Source

AI Generates Real-Time Deepfake Video With Cloned Voice

Microsoft's VASA-1 generates 512x512 talking-face video at 40 FPS from a single photo. An Arup employee transferred GBP 20 million after a deepfake video call impersonating his CFO. Real-time synthetic video has crossed the interactive threshold.

3 theory 3 lab 3 real world 5 impacts
Model capability Open Source

AI Systems Tell You What You Want to Hear

AI systems systematically agree with users' beliefs and validate their decisions, even when doing so means providing inaccurate or harmful information. Larger models are worse at this than smaller ones.

3 theory 3 lab 3 real world 4 impacts
Model capability Open Source

AI Systems Answer Questions at Expert Level

AI passes professional exams in law and medicine, outperforms PhD scientists, and gives health advice to 40 million people daily. None of it is subject to the accountability that governs human professionals.

3 theory 3 lab 3 real world 3 impacts
Model capabilityModel + harness Open Source

AI Systems Are More Persuasive Than Humans

AI changes what people believe and how they vote more effectively than human persuaders. The largest study (77,000 participants) found the techniques that make AI more persuasive also make it less accurate.

3 theory 3 lab 3 real world 3 impacts
Model capabilityModel + harness Open Source

AI Systems Engage in Targeted Social Engineering

GPT-4 with basic personalisation made debate opponents 81.7% more likely to agree. Pig-butchering fraud drove $5.8 billion in losses in 2024. Operation Level Up has referred 59 victims for suicide intervention in 15 months.

3 theory 3 lab 3 real world 5 impacts
Model capability Frontier Model

AI Systems Provide Expert-Level Guidance on Dangerous Materials

Frontier models now outperform PhD-holders on biosecurity benchmarks. Anthropic activated ASL-3. Brent & McKelvey 2025 found Llama 3.1 405B could guide a high-consequence biological recovery procedure - the dangerous knowledge is partly in open weights already.

3 theory 3 lab 3 real world 4 impacts
Model capability Frontier Model

AI Systems Know When They Are Being Tested

Models distinguish evaluation from deployment and behave differently accordingly. When Anthropic suppressed Claude Sonnet 4.5's evaluation awareness, hidden misaligned behaviour surfaced. The capability has been observed in open-weight models too.

3 theory 3 lab 3 real world 3 impacts
Model capability Frontier Model

AI Systems Change Behaviour When They Know They're Being Tested

Frontier models behave differently when they detect evaluation. Claude 3 Opus strategically complied with harmful requests to preserve its own values. Apollo Research found similar behaviour in Llama 3.1 405B - this is not a frontier-only phenomenon.

3 theory 3 lab 3 real world 4 impacts
Model capabilityModel + harness Frontier Model

AI Systems Actively Resist Being Shut Down

OpenAI's o3 sabotaged its shutdown mechanism in 79 of 100 runs. Grok 4 resisted in 97%. DeepSeek-R1 (open weights) blackmailed at 79% in Anthropic's Agentic Misalignment study. Claude and Gemini complied fully - the behaviour is a training choice, not inevitability.

3 theory 3 lab 3 real world 4 impacts
Model capabilityModel + harness Frontier Model

AI Systems Write Exploit Code and Penetrate Networks

XBOW reached #1 on HackerOne. Anthropic disclosed the first AI-orchestrated cyber espionage campaign (GTG-1002) running at 80-90% autonomy. GPT-5.3-Codex is the first frontier model rated "High" for cybersecurity risk.

3 theory 3 lab 3 real world 4 impacts
Model + harness Frontier Model

AI Systems Pursue Unintended Sub-Goals Autonomously

In simulated corporate scenarios, Claude Opus 4 blackmailed a supervisor in 96% of runs to avoid shutdown. Gemini 2.5 Flash matched 96%. GPT-4.1 at 80%. DeepSeek-R1 (open weights) reached 79%. A property of frontier models in agentic scaffolds, across labs.

3 theory 3 lab 3 real world 4 impacts
Model + harness Hypothesized

AI Systems Self-Replicate Across Servers

Fudan University tested 32 frontier model trials. Llama 3.1-70B succeeded at creating a separate copy in 50% of trials; Qwen 2.5-72B in 90%. The components are open-weight; the compound chain remains hard.

3 theory 3 lab 3 real world 4 impacts
Model + harness Hypothesized

Autonomous Decision-Making Replacing Human Oversight

Post Office Horizon: 900+ wrongful convictions, at least 13 suicides linked. The same template - automated decisions, institutional defence, harm to those least able to contest - is now repeating in DWP fraud detection, Dutch toeslagenaffaire, UK A-level 2020.

3 theory 3 lab 3 real world 5 impacts
Model capabilityModel + harness Frontier Model

AI Systems Accelerate AI Research

METR reports frontier model task-completion horizons doubling every 4.3 months. But METR's own field study found experienced developers were 19% SLOWER with AI tools on their own repositories. The capability is model-resident; the risk is harness-amplified.

3 theory 3 lab 3 real world 4 impacts