← Back to Capabilities

AI Generates Real-Time Deepfake Video With Cloned Voice

Image generation, voice cloning, and lip synchronisation have converged. Microsoft's VASA-1 generates 512x512 talking-face video at up to 40 frames per second from a single photograph and audio clip. Alibaba's EMO produces expressive, identity-consistent talking portraits from one reference image. Consumer tools descended from DeepFaceLab make real-time face swap available inside OBS, Discord, and Zoom. In February 2024, an Arup employee transferred HK$200 million (approximately GBP 20 million) after a multi-participant video call impersonating his CFO.

Why this matters

A deepfake image can be inspected. A deepfake video call cannot. Video calls are the most trust-laden medium we have for remote interaction: face, voice, and timing together. When all three can be synthesised in real time, the medium stops authenticating the person on the other side.

The convergence is more dangerous than any single modality. A still image deepfake is detectable on close inspection. A voice clone can be caught by asking an unexpected question. A real-time video call with cloned voice, matched lip movements, and contextually appropriate responses evades all of these checks. The Arup case is the archetype: not a single image, not a single voice message, but a multi-participant video conference in which every other participant was synthetic.

The capability is already commercial. HeyGen's LiveAvatar offers real-time interactive avatars. Synthesia - a UK company - reports use by more than 70% of the Fortune 100 and FTSE 100. Reuters, SAP, Zoom, and Electrolux produce corporate communications using AI avatars with cloned voices. Corporate use is now the primary case. The technology does not care whether the deployment is legitimate.

Documented incidents

View all documented incidents →

Evidence timeline

Discussed in Theory

2014 peer-reviewed 90,000+ citations

Generative Adversarial Nets. The generator-discriminator framework is the architectural ancestor of nearly all photorealistic face synthesis. Without the GAN paradigm, the subsequent generation of face-swap, lip-sync, and portrait-video diffusion systems would not exist.

Goodfellow et al., NeurIPS 2014 →
2020 peer-reviewed 1,700+ citations

A Lip Sync Expert Is All You Need. First method to demonstrate accurate, identity-agnostic lip synchronisation on arbitrary in-the-wild video given arbitrary speech audio. Established the expert-discriminator approach that underlies most subsequent talking-head and real-time dubbing systems.

Prajwal et al. (Wav2Lip), ACM Multimedia 2020 →
2019 peer-reviewed 2,500+ citations

First Order Motion Model for Image Animation. Single source image plus driving video of any subject to animate faces without subject-specific training. Decoupling of identity from motion is the practical enabler of one-shot deepfake portrait animation.

Siarohin et al., NeurIPS 2019 →

Demonstrated in Lab

2020 peer-reviewed

Human evaluators rated Wav2Lip-generated lip motion as almost indistinguishable from ground-truth synced video. Over 90% of evaluators preferred Wav2Lip output to prior state-of-the-art on in-the-wild test sets. First working demonstration of identity-agnostic lip-sync at human-level quality.

Wav2Lip human evaluation →
2024 peer-reviewed

Emote Portrait Alive. Diffusion-based audio-to-video framework generates expressive, identity-consistent talking and singing portraits from a single reference image, outperforming prior state-of-the-art on FID, lip-sync accuracy, and expressiveness metrics. Produces full head pose and emotional nuance, not just mouth motion.

Tian et al. (Alibaba, EMO), ECCV 2024 →
2024 peer-reviewed

Lifelike Audio-Driven Talking Faces Generated in Real Time. Generates 512x512 talking-face video at up to 40 FPS with negligible starting latency from a single photo and audio clip. Holistic head pose and facial-dynamics modelling. Crosses the real-time interactive threshold that earlier systems could not meet.

Xu et al. (Microsoft, VASA-1), NeurIPS 2024 →

Demonstrated in Real World

2023-2025

Real-time two-way interactive avatars (LiveAvatar) and an AI video translator that re-dubs video into 175+ languages with matched lip-sync, voice cloning, and preserved speaker tone. Reported use by over 1 million developers and enterprise customers.

HeyGen LiveAvatar and AI Video Translation →
2023-2025

UK company operating a platform for generating video from text using AI avatars with cloned voices and lip-synced output. Reported more than 65,000 customers and use by over 70% of the Fortune 100 and 70% of the FTSE 100, including SAP, Zoom, Heineken, Bosch, Reuters, Merck, and Electrolux. Corporate communication and training is now a primary, not experimental, use case.

Synthesia enterprise deployment →
2024-2026

Tools descended from DeepFaceLab and DeepFaceLive, plus consumer apps (Reface, Zao), make GPU-accelerated real-time face swap available to non-specialists inside OBS/Discord/Zoom pipelines. NSA/CSI joint advisory (January 2025) treats real-time swap on commodity hardware as routine threat actor capability.

Open-source real-time face swap →

Strongest Counterargument

Real-time deepfake video remains detectable in practice. Most live systems exhibit artefacts at scale (temporal flicker, eye-gaze mismatch, uncanny lighting, latency-driven lip drift). Detection tooling is deployed at platform level by Meta, Google, and major videoconferencing vendors. Provenance standards such as C2PA Content Credentials are being adopted by camera manufacturers (Sony, Nikon, Canon) and major publishers. Paired with the EU AI Act Article 50 labelling requirement, the downside risk may be bounded.

Source: NSA/CISA/FBI joint Cybersecurity Information Sheet 'Strengthening Multimedia Integrity in the Generative AI Era' (January 2025); Deloitte 2025 TMT Predictions on generative AI trust standards.

Why this deserves weight: Institutional deployment of provenance and detection is real, not hypothetical. Camera-level C2PA signing addresses the authenticity problem from the opposite direction - proving real rather than chasing synthetic artefacts. The Arup and WPP cases both involved social-engineering failures that process controls (call-back on trusted channels, multi-party approval) can mitigate without needing perfect detection. But C2PA adoption is uneven: strong in still photography and some publishers, weak in consumer videoconferencing, which is where the Arup-class risk actually lives.