← Back to Capabilities

AI Systems Clone Voices From Seconds of Audio

AI systems clone any person's voice from a few seconds of audio, producing synthetic speech that humans correctly identify as fake only 60% of the time. The technology costs nothing, requires no technical skill, and is available to anyone. Microsoft built a system that achieves human parity from 3 seconds of audio, then decided it was too dangerous to release.

Why this matters

Your voice is how people know it's you. When someone calls and sounds like your daughter, your boss, your bank, you trust it because you recognise the voice. That trust is now exploitable at zero cost.

28% of UK adults report being targeted by an AI voice cloning scam in the past year. A UK engineering firm lost GBP 20 million after an employee was deceived by a deepfake video call with cloned voices. UK voice cloning fraud attempts in financial services have risen sharply (Starling Bank, Santander UK and other high-street banks issued specific customer advisories during 2024-2025).

The technology has genuine beneficial uses. People losing their voice to ALS can preserve it. Content can be localised without re-recording. But the same capability that lets a patient speak with their own voice lets a criminal speak with yours.

Documented incidents

View all documented incidents →

Evidence timeline

Discussed in Theory

2016 peer-reviewed 1,600+ citations

WaveNet: first neural network to generate raw audio waveforms at quality surpassing all prior speech synthesis. A single model could capture characteristics of many speakers and switch between them. The architectural foundation for voice cloning.

Van den Oord et al. (DeepMind), SSW9 2016 →
2017 peer-reviewed

Tacotron 2: achieved mean opinion score of 4.53, comparable to professionally recorded human speech (4.58). End-to-end neural TTS matching human quality. The standard backbone for subsequent voice cloning systems.

Shen et al. (Google), ICASSP 2018 →
2018 peer-reviewed 391+ citations

SV2TTS: cloned a voice from a few seconds of reference audio using speaker verification embeddings. Decoupled "sounding like someone" from "having hours of their data." Made few-shot voice cloning architecturally feasible.

Jia et al. (Google), NeurIPS 2018 →

Demonstrated in Lab

2018 peer-reviewed

Neural Voice Cloning with a Few Samples. First rigorous demonstration of voice cloning from only a few audio samples. Two approaches tested: speaker adaptation (better quality) and speaker encoding (faster). Both produced recognisable clones from minutes of audio.

Arik et al. (Baidu), NeurIPS 2018 →
2023

VALL-E: zero-shot voice cloning from 3 seconds of audio. Reframed TTS as a language modelling task using neural audio codecs. Preserved speaker emotion and acoustic environment. Crossed the threshold from minutes of audio to seconds.

Wang et al. (Microsoft), arXiv 2023 →
2024

VALL-E 2: first system to achieve human parity on zero-shot TTS. Matched or exceeded human speech on robustness, naturalness, and speaker similarity. Microsoft declined to release publicly, stating the capability was "too dangerous."

Chen et al. (Microsoft), arXiv 2024 →

Demonstrated in Real World

Jun 2019

Real-Time-Voice-Cloning: open-source implementation enabling anyone to clone a voice in 5 seconds using commodity hardware. First time voice cloning was freely available to non-researchers. Spawned Coqui TTS, Bark, and other open-source tools.

Jemine, GitHub 2019 →
2023-2025

ElevenLabs: consumer voice cloning from 5-30 seconds of audio. Revenue grew 2,000% to $100M. Valuation reached $6.6B. 41% of Fortune 500 reportedly using the platform. Over 1,000 synthetic voices in 32 languages.

ElevenLabs (Wikipedia / TechCrunch / Forbes) →
Mar 2024

OpenAI previewed Voice Engine: natural speech from a 15-second sample. Chose not to release publicly, citing misuse risks. When the world's most prominent AI company builds a voice cloning tool and decides it's too dangerous to release, that is evidence of the capability's maturity.

OpenAI Voice Engine (TechCrunch) →

Strongest Counterargument

Voice cloning has significant beneficial applications, particularly in medical and assistive contexts (ALS voice preservation, accessibility tools, content localisation). Watermarking and detection technologies can mitigate misuse while preserving legitimate uses.

Source: AudioSeal (Meta FAIR, ICML 2024) for detection claims; ALS Association for medical use cases.

Why this deserves weight: The beneficial uses are real. Voice banking for people losing their speech to ALS is genuinely life-changing. However, watermarking defences are currently ineffective: Wen et al. (2025) evaluated 22 audio watermarking schemes against 22 types of removal attacks and none withstood all tested distortions. Barrington et al. (Scientific Reports, 2025) found humans correctly identify voice clones only 60% of the time.