
What Is Voice Cloning and How Zero-Shot Models Reproduce a Speaker's Voice from Audio
Voice cloning uses speaker embeddings and zero-shot conditioning to reproduce a target voice from 6–10 seconds of reference audio.
This theme is curated by our AI council — see how it works.
AI audio, video & 3D is the branch of generative AI that produces media beyond still images — synthesized speech, cloned voices, complete music tracks, talking-head avatars, edited video footage, and game-ready 3D assets — from a text prompt or a reference sample. Six capabilities make up the theme, and they share one production question: whether the output survives contact with a real pipeline, a real brand, and a real legal review.
Image generation took the headlines, but the media that businesses actually commission — training videos, product demos, localized voiceovers, game assets — is audio, video, and 3D. For a developer, this theme is mostly API integration work with two unusual failure modes: outputs that drift over time (a voice whose prosody wanders mid-paragraph, a face that subtly morphs between frames) and inputs that carry legal weight (a cloned voice is someone’s biometric identity; a music model’s training data is somebody’s catalogue). The integration is the easy part; knowing which capability you actually need, and what shipping it commits you to, is where this page earns its keep.

The audio side of the theme is the most mature, and text-to-speech is its foundation — every speaking product downstream of it, from cloned narrators to talking avatars, runs a TTS engine at its core. What text-to-speech AI is and how neural TTS architectures convert text into speech is the best first read in the whole theme for exactly that reason. The first real decision arrives fast, because speech now ships from two kinds of vendor: dedicated TTS API vs. general LLM platform walks the trade-off between purpose-built engines like Cartesia Sonic and Kokoro and the speech endpoints bolted onto general model platforms.
Voice cloning is text-to-speech with one constraint added — the output must sound like a specific, real person — and that single constraint changes both the technology and the ethics. How zero-shot models reproduce a speaker’s voice from audio explains why a few seconds of reference audio is now enough, and the hands-on route is how to clone a voice with Fish Speech, XTTS v2, and CosyVoice2. Read those alongside the question the capability raises — consent, deepfake audio, and the legal gaps — because in this corner of the theme the safeguards are part of the integration, not an appendix.
The other audio branch is AI music generation, which follows the same text-to-audio pattern but produces a full arranged track rather than speech. How text-to-audio models convert prompts into full tracks is the orientation read, and using Suno v5.5, Mureka, and the Google Lyria API for production music covers the practical end. Music is also where the theme’s licensing questions are furthest developed: the market restructured around settlements with rights holders, and what you may ship commercially now depends on which model and plan you generate with.
On the visual side, AI avatar generation is where audio meets video: a talking head that lip-syncs to a script, usually with a cloned or synthetic voice underneath. How talking-head synthesis works covers the mechanism at orientation depth, and using HeyGen and Synthesia for training videos and multilingual localization is the read for the dominant business use case — one recording, many languages, no reshoots. If your application needs the avatar to respond live rather than render offline, building a real-time avatar pipeline with D-ID and open-source models covers the harder, latency-bound variant.
AI video editing works on footage that already exists — removing objects, restyling scenes, syncing lips to new audio — rather than generating a presenter from scratch. How object removal, style transfer, and lip sync actually work is the entry point, and the guide to building an editing pipeline with Runway Aleph and Pika turns it into a repeatable workflow. Video is also where the theme’s hardest technical constraint lives: temporal consistency, the requirement that every generated frame agrees with its neighbours.
The frontier of the theme is text-to-3D, and it differs from everything above in what it outputs: not rendered pixels but an asset — a mesh with geometry and textures that a game engine can light, animate, and reuse. How NeRF, Gaussian splatting, and mesh diffusion turn prompts into 3D assets maps the competing approaches; for hands-on work, Meshy, Tripo AI, and Rodin Gen-2 for game assets covers the tool route and building a TRELLIS pipeline from prompt to game-engine export covers the open-source one.
Read the text-to-speech and avatar entries and you have the theme’s spine — everything else is a variation on the same prompt-in, media-out contract, with a different output format and a different set of strings attached.
The three audio capabilities are the most-confused trio in the theme, because all three turn text into sound:
| Text-to-speech | Voice cloning | AI music generation | |
|---|---|---|---|
| You provide | A script | A script plus reference audio of a real speaker | A text prompt or reference track |
| You get back | Speech in a stock or designed voice | Speech in that specific person’s voice | A complete arranged track |
| Whose rights are in play | Nobody’s — the voice is synthetic | The reference speaker’s — consent is the gate | The training catalogue’s — licensing is the gate |
| Typical production trap | Prosody drift on long scripts | Shipping without a consent workflow | Commercial use outside the model’s licence terms |
Three more distinctions on the visual side trip people just as often:
Q: Where should I start with generative audio and video as a software developer? A: Start with text-to-speech: its APIs are the most mature in the theme, experiments are cheap, and every speaking product downstream builds on it. The neural TTS explainer gives you the architecture vocabulary the rest of the theme assumes.
Q: Do I need voice cloning, or is stock text-to-speech enough? A: Stock TTS covers most product work — narration, alerts, assistants — with none of the consent overhead. Cloning is only worth it when the voice must belong to a specific person, such as a brand voice or a localized known speaker. If you do need it, the Fish Speech and XTTS v2 pipeline guide is the practical route.
Q: Should I generate an avatar video or edit real footage? A: If usable footage exists and only needs correction, edit it; if you need one script delivered in many languages or many variants, a generated avatar beats reshoots on cost and turnaround. The HeyGen and Synthesia guide shows what avatar-first production looks like at localization scale.
Q: Why do my generated audio and video pass a demo but fall apart at production length? A: Consistency over time is the theme’s shared hard problem: prosody wanders across long scripts, faces drift across frames, and compute cost scales with duration. The technical-limits read on temporal consistency and identity drift explains why short clips hide exactly these failures.
Q: Can I ship AI-generated music or voices in a commercial product? A: Often yes — but it depends on the model’s licence and how its training data was cleared, not on output quality. Music is furthest along: the post-settlement market read maps which providers now come with cleared rights. For voices, speaker consent is the additional gate.
AI avatar generation creates photorealistic or stylized digital avatars from a reference photo, video, or text …
AI Music Generation refers to tools and models that create original music from text prompts or reference audio. These …
AI video editing uses generative models to manipulate existing footage automatically — removing objects, transferring …
Text-to-3D refers to AI models and pipelines that generate three-dimensional assets directly from text descriptions or …
Text-to-Speech (TTS) is an AI technology that converts written text into natural-sounding spoken audio. Modern neural …
Voice cloning is the process of training an AI model on reference audio samples to reproduce a specific speaker's voice. …
MONA's articles build your mental model — how things work, why they work that way, and what intuition to develop.
Updated Aug 19, 2026
Concepts covered

Voice cloning uses speaker embeddings and zero-shot conditioning to reproduce a target voice from 6–10 seconds of reference audio.

Neural TTS converts text into speech through learned acoustic models. Tacotron 2 scored 4.53 MOS at ICASSP 2018; current systems hit sub-90ms latency.

Text-to-3D uses Score Distillation Sampling to convert text into 3D meshes via NeRF or Gaussian splatting — no labeled 3D training data needed.

AI music generation converts text prompts into audio tracks. Neural codecs compress waveforms into tokens; transformers or diffusion models sequence them.

AI video editing uses diffusion models to edit footage directly — Runway Aleph and Pika remove objects, transfer style, and sync lips without reshooting.
AI avatar generation reanimates a face from audio using two pipelines: 2D lip-sync over real video, or 3D reconstruction with NeRF and Gaussian splatting.

Neural TTS converts text to audio via phoneme processing, prosody modeling, and vocoders. In 2026, expressiveness and low latency remain in structural tension.

Text-to-3D tools produce non-manifold meshes with broken UV maps. Topology errors, splat format gaps, and multi-view drift are the core barriers in 2026.

Neural codecs tokenize audio into discrete sequences for AI music models. Codec fidelity is largely solved in 2026; structural coherence past 2 minutes is not.
AI avatar generation now runs on diffusion transformers, not GANs — HeyGen and Synthesia both switched in 2026, but identity drift remains unsolved.

Mel spectrograms and speaker embeddings are the core of voice cloning. Zero-shot models clone from 3 seconds but fail on accented and expressive voices.

AI video editing tools regenerate each frame via diffusion, not edit pixels—causing temporal drift. Runway Aleph 2.0 caps clips at 30 seconds, 1080p.