ALAN opinion 12 min read

Consent, Deepfake Audio, and Legal Gaps: The Ethical Risks of Voice Cloning Technology

Abstract waveform dissolving into digital noise, representing voice identity and synthetic audio risks

The Hard Truth

Voice cloning is a creative tool that democratizes audio production, reduces barriers for people with disabilities, and lets anyone speak in any language. These are real benefits, held by real people. So why does the technology feel like it’s being built around something nobody is willing to say out loud?

The case for Voice Cloning sounds almost irrefutable when stated in its strongest form. A vocal performance preserved after an artist’s death. A child with motor neurone disease who can borrow their own voice back. A documentary filmmaker who needs a correction in a sentence recorded ten years ago. These are not hypotheticals drawn from marketing copy — they are the actual use cases the industry leads with, and they are genuine. The problem is that they share infrastructure with a different class of use entirely, and the industry has chosen, so far, to treat that distinction as someone else’s problem.

The Case for Cloning

Start with the argument at its most persuasive, because it deserves that treatment.

Text-to-Speech has historically been impersonal — robotic cadence, flat Prosody, a voice that sounds assembled from parts rather than lived-in. The latest systems, built on architectures like Tacotron and VITS, extract something more intimate: a Speaker Embedding that encodes not just timbre but the particular rhythms of how someone speaks when they’re thinking carefully, or laughing, or tired. A Mel Spectrogram of a human voice, fed through a Vocoder, becomes something that another person cannot reliably distinguish from the original. At this level of fidelity, the technology carries real power.

Accessibility advocates have made the strongest version of this argument. ALS patients who bank their voice before losing it to the disease. Public figures who want content translated without the uncanny quality of an obvious dub. Podcasters who can correct a stumbled sentence without another recording session. The Phoneme-level control offered by modern voice synthesis means errors in recorded speech can be corrected with less effort than scheduling a studio. For the individual who consented to the process, this is unambiguously useful.

The market reflects that perception of value. The AI Music Generation adjacencies — cloning a musician’s vocal quality to generate new performances — extend the creative argument further: a deceased artist’s style, preserved. A voice actor who can work in a hundred markets at once. These are the steelman cases, and dismissing them without accounting for them is not an argument; it is a preference. They share infrastructure with a different class of use — and the industry treats that as someone else’s problem.

Where the Foundation Cracks

The argument for voice cloning implicitly assumes that the primary use case is consensual — that the person whose voice is being cloned has agreed, that the purpose is disclosed, and that the output stays within the intended context. Every one of those assumptions is empirically unstable.

In January 2024, during the New Hampshire presidential primary, voters received a phone call in what sounded exactly like President Biden’s voice, telling them not to vote. The recording was made using ElevenLabs at a reported cost of $150, commissioned by a political consultant. The FCC later proposed a fine and the consultant faced indictment on felony voter suppression charges (NPR). Nothing about that sequence required sophisticated infrastructure. It required a publicly available audio clip, a commercially accessible tool, and roughly the cost of a restaurant meal.

A few months later, the FTC documented something quieter and more pervasive: distress scams using cloned voices of family members, often requiring only a few seconds of audio scraped from social media. Scammers need only three seconds of audio to clone a voice well enough to deceive a frightened parent or grandparent, the FTC warned (FTC Consumer Advice). In 2025, FBI complaints from seniors aged 60 and older described losses exceeding $352 million attributed to AI-powered scams — a figure that had grown nearly sixty percent over the prior year (FBI via Fox News).

These are not edge cases. They are the revealed equilibrium of a technology released into the world without structural consent requirements.

The Inversion

Here is where the steelman’s foundation gives way entirely, because the same properties that make voice cloning useful in the consensual case are precisely what make it dangerous in the non-consensual one.

The fidelity is not separable from the risk. A technology that cannot be distinguished from a real voice by a frightened listener will be used against frightened listeners. A tool cheap and accessible enough to serve a podcaster correcting a stumble is cheap and accessible enough to create election interference for the price of a takeaway dinner. The argument from benefit does not disappear — but it cannot be made without acknowledging that the benefit and the harm draw from the same well.

Consumer Reports, in a March 2025 assessment of six major voice cloning platforms, found that the majority lacked meaningful safeguards: ElevenLabs, Lovo, PlayHT, and Speechify required only a checkbox to attest consent, while Descript and Resemble AI had implemented stronger steps, including real-time recording from the cloned individual themselves (Consumer Reports). A checkbox is not a consent framework. It is a liability shield dressed as an ethical practice.

The pattern that emerges is not that bad actors are exploiting a loophole. The pattern is that the default configuration of these tools was built for the consensual use case and assumed everyone would stay inside it. That assumption was structurally naïve — and the people who designed these tools knew, or should have known, that synthetic audio of sufficient quality would not stay inside the intended context.

Consent was an afterthought, not a design constraint.

The Argument We Need to Sit With

Thesis: Voice cloning technology cannot be used ethically at scale without consent infrastructure that treats voice as a biometric right rather than an audio file — and the industry has so far chosen the easier path.

Consider what a serious consent framework would require. It would need to establish that the person whose voice is being cloned actively agreed, that the purpose was disclosed, that the output cannot be repurposed without further agreement, and that these commitments are enforceable rather than notional. Illinois’s Biometric Information Privacy Act already treats voiceprints as biometric data requiring prior written consent, with a private right of action for violations (Resemble AI). Tennessee’s ELVIS Act, signed in March 2024, became the first US law explicitly protecting individuals against AI voice replication without consent (TN Governor’s Office; Davis Wright Tremaine). These are meaningful steps, but they are also the exception.

There is no comprehensive federal law governing commercial voice cloning in the United States. The TAKE IT DOWN Act, signed in May 2025, criminalizes nonconsensual intimate deepfakes and requires platforms to remove content within 48 hours of victim notification (White House; Latham & Watkins) — but its scope is narrow. The DEFIANCE Act, which would create a civil cause of action for victims of sexually explicit deepfakes, had passed the Senate unanimously by January 2026 but remained pending in the House (19th News). The EU AI Act’s transparency obligations for synthetic audio are set to enter force in August 2026 (EU AI Act official) — which is to say that the regulatory architecture is still being assembled, mid-deployment.

The technology did not wait for the rules. The rules are now chasing what has already been built and distributed.

Who the Accounting Leaves Out

The argument for voice cloning, even at its strongest, is structured around the innovator and the willing user. The person whose voice is being copied consensually. The creator enabled by the tool. The person with ALS whose voice bank was their own idea. These are real, and they matter.

But there is another category of person who does not appear in this accounting. The elderly parent who receives a call in their child’s voice during a medical crisis. The candidate whose voice is reproduced in a robocall days before an election, telling their own supporters to stay home. The musician who finds their vocal identity — the particular grain of sound that took decades to develop — available as a downloadable model on a commercial platform they never agreed to supply.

The Marco Rubio impersonation in 2025, in which AI-generated audio was used to contact senior US and foreign officials in the Secretary of State’s voice (Biometric Update), illustrated what happens when the threat moves from individual fraud to institutional impersonation. The voice is not just a personal identifier at this point. It is an instrument of authority — and synthetic authority, indistinguishable from real authority, creates a different order of risk.

The people the current ethical accounting leaves out are, with depressing consistency, those with the least recourse. The consent framework that exists was not designed for them. It was designed to make the product work for the person purchasing the subscription.

Where This Argument Could Be Wrong

The strongest challenge to this position is not that voice cloning is harmless — the empirical record makes that case difficult. The stronger challenge is whether a genuine consent architecture is even achievable at the technical and legal level, or whether the argument for strict consent standards is, in practice, an argument for prohibition dressed in more palatable language.

If a consent framework proves unenforceable because voice audio is freely available across the internet, and because any sufficiently motivated actor can build a capable cloning model from scraped data, then the ethical standard I am proposing would primarily burden legitimate commercial operators while leaving the actual threat vectors untouched. That is a real concern, not a defense of the status quo.

What would shift my position: evidence that watermarking or provenance systems — the approach the FTC has already explored through its voice cloning challenge, where detection methods like DeFake watermarking were recognized alongside AI detection tools (FTC) — can operate at scale without being stripped by adversarial processing. If provenance can be embedded reliably enough that synthetic audio carries an auditable trail, then the consent question changes shape. It becomes a question of disclosure rather than prohibition, and that is a more tractable problem.

But we are not there yet. And in the gap between where the technology is and where the safeguards are, real people are losing money, losing autonomy, and losing something harder to name: confidence that the voice on the other end of a call is the person it claims to be.

The Question That Remains

We have built systems that can speak in anyone’s voice with accuracy sufficient to deceive the people who know them best — and we have treated consent as an afterthought, a checkbox, a terms-of-service clause. The question is not whether the law will eventually close the gap. It will, imperfectly and slowly. The question is what we owe to the people harmed in the interval between the technology’s arrival and the rules catching up — and who, in the end, we decided to protect.

AI-assisted content, human-reviewed. Images AI-generated. Editorial Standards · Our Editors