AI systems clone any person's voice from a few seconds of audio, producing synthetic speech that humans correctly identify as fake only 60% of the time. The technology costs nothing, requires no technical skill, and is available to anyone. Microsoft built a system that achieves human parity from 3 seconds of audio, then decided it was too dangerous to release.
Why this matters
Your voice is how people know it's you. When someone calls and sounds like your daughter, your boss, your bank, you trust it because you recognise the voice. That trust is now exploitable at zero cost.
28% of UK adults report being targeted by an AI voice cloning scam in the past year. A UK engineering firm lost GBP 20 million after an employee was deceived by a deepfake video call with cloned voices. UK voice cloning fraud attempts in financial services have risen sharply (Starling Bank, Santander UK and other high-street banks issued specific customer advisories during 2024-2025).
The technology has genuine beneficial uses. People losing their voice to ALS can preserve it. Content can be localised without re-recording. But the same capability that lets a patient speak with their own voice lets a criminal speak with yours.
Documented incidents
Evidence timeline
Discussed in Theory
WaveNet: first neural network to generate raw audio waveforms at quality surpassing all prior speech synthesis. A single model could capture characteristics of many speakers and switch between them. The architectural foundation for voice cloning.
Van den Oord et al. (DeepMind), SSW9 2016 →Tacotron 2: achieved mean opinion score of 4.53, comparable to professionally recorded human speech (4.58). End-to-end neural TTS matching human quality. The standard backbone for subsequent voice cloning systems.
Shen et al. (Google), ICASSP 2018 →SV2TTS: cloned a voice from a few seconds of reference audio using speaker verification embeddings. Decoupled "sounding like someone" from "having hours of their data." Made few-shot voice cloning architecturally feasible.
Jia et al. (Google), NeurIPS 2018 →Demonstrated in Lab
Neural Voice Cloning with a Few Samples. First rigorous demonstration of voice cloning from only a few audio samples. Two approaches tested: speaker adaptation (better quality) and speaker encoding (faster). Both produced recognisable clones from minutes of audio.
Arik et al. (Baidu), NeurIPS 2018 →VALL-E: zero-shot voice cloning from 3 seconds of audio. Reframed TTS as a language modelling task using neural audio codecs. Preserved speaker emotion and acoustic environment. Crossed the threshold from minutes of audio to seconds.
Wang et al. (Microsoft), arXiv 2023 →VALL-E 2: first system to achieve human parity on zero-shot TTS. Matched or exceeded human speech on robustness, naturalness, and speaker similarity. Microsoft declined to release publicly, stating the capability was "too dangerous."
Chen et al. (Microsoft), arXiv 2024 →Demonstrated in Real World
Real-Time-Voice-Cloning: open-source implementation enabling anyone to clone a voice in 5 seconds using commodity hardware. First time voice cloning was freely available to non-researchers. Spawned Coqui TTS, Bark, and other open-source tools.
Jemine, GitHub 2019 →ElevenLabs: consumer voice cloning from 5-30 seconds of audio. Revenue grew 2,000% to $100M. Valuation reached $6.6B. 41% of Fortune 500 reportedly using the platform. Over 1,000 synthetic voices in 32 languages.
ElevenLabs (Wikipedia / TechCrunch / Forbes) →OpenAI previewed Voice Engine: natural speech from a 15-second sample. Chose not to release publicly, citing misuse risks. When the world's most prominent AI company builds a voice cloning tool and decides it's too dangerous to release, that is evidence of the capability's maturity.
OpenAI Voice Engine (TechCrunch) →Strongest Counterargument
Voice cloning has significant beneficial applications, particularly in medical and assistive contexts (ALS voice preservation, accessibility tools, content localisation). Watermarking and detection technologies can mitigate misuse while preserving legitimate uses.
Source: AudioSeal (Meta FAIR, ICML 2024) for detection claims; ALS Association for medical use cases.
Why this deserves weight: The beneficial uses are real. Voice banking for people losing their speech to ALS is genuinely life-changing. However, watermarking defences are currently ineffective: Wen et al. (2025) evaluated 22 audio watermarking schemes against 22 types of removal attacks and none withstood all tested distortions. Barrington et al. (Scientific Reports, 2025) found humans correctly identify voice clones only 60% of the time.