VSL·TOOLS

BLOG Blog

AI Voice Cloning & Lip Sync for Video Ads: How It Works and Cost

·7 min read

Two tools have changed video creative production more than anything else: AI voice cloning and lip sync. The first removes the voice actor and the recording studio; the second removes the shoot. Together they make an AI spokesperson video possible in which a real person says any text in any language — from one photo and ten seconds of audio.

Let's break down how it works, what drives quality and what a minute of such video costs.

AI voice cloning: how the model learns a voice

A voice clone is a speech synthesis model that was shown a sample and asked to speak "the same way." Modern models need just 10 seconds of clean speech: they don't "memorize" the words, they extract characteristics — timbre, pitch, pace, manner, breathing — and reproduce them on new text.

What happens inside:

  1. Transcribing the sample. The service recognizes what's said in the recording and links sound to text — that's how the model learns how this voice pronounces specific sounds.
  2. Extracting the voice "fingerprint." A compact, text-independent description of the voice is pulled from the recording.
  3. Synthesis. New text is voiced with that fingerprint: intonation is built from punctuation and the meaning of the phrase.

The clone can speak languages that weren't in the sample: an English voice reads Spanish or German text without trouble, keeping its timbre. The accent, though, is inherited from the original — for local geos, use a sample from a native speaker.

What makes a voice clone for ads sound right

  • A clean sample. No music, reverb or second voice. A voice message from a phone in a quiet room beats a studio recording with effects.
  • Natural speech. A sample "as in conversation," not reading from a page — otherwise the clone will read.
  • Short lines. The model holds intonation for one or two sentences; split a long paragraph.
  • Control over stress and pauses. In names and rare words the stress is set by hand; pauses — with punctuation. A good service keeps a database of problem words and inserts the marks itself.

A voice designed from a description

When there's no sample, or someone else's voice can't be used, the voice is designed: you set gender, age, timbre and mood — and get a unique voice not tied to any real person. For brands it's also the legally clean option: nobody will show up claiming "that's my voice."

AI lip sync: lips on every word

Lip sync is the generation of mouth movement and facial expressions to match a specific audio track. Input: a photo or video of the character plus audio; output: a lip sync AI video in which the character says exactly this text.

Two approaches:

  • Lip sync on video. You have a finished video with a spokesperson and need to replace the speech: the AI redraws the mouth area to fit the new track, the rest of the frame stays original. That's how finished videos get translated and re-voiced.
  • Generation from a photo. You only have a photograph: the model creates the whole video — head movement, facial expressions, blinking, gestures — and syncs the lips to the voiceover. This is what's usually called a "talking head" or "animating a photo," but modern models change shots and angles as if the video had been filmed with several cameras.

What separates good lip sync from a "mask"

  • Consonants read. "P," "B" and "M" need the lips to close; if the mouth is open on those sounds, the viewer senses something's off, even without knowing why.
  • Facial expressions tied to speech. Eyebrows, cheeks and head tilt move with the intonation instead of living on their own.
  • Natural pauses. Between lines the mouth closes and the pose doesn't "jump."
  • A stable face. Features don't change from scene to scene; in close-ups the skin and teeth are detailed.

Different models balance speed and detail differently, which is why the VSL·TOOLS studio has three tiers: Lite — fast and cheap for tests, Standard — natural gestures and facial expressions, Pro — maximum facial detail for close-ups. The service can pick the model itself by the length of the video's segments.

How voice clone + lip sync come together in a creative

  1. The text is split into lines.
  2. Each line is voiced by the clone — as a separate file, so you regenerate a phrase, not the video.
  3. The character from the photo speaks the lines; scenes change between lines.
  4. Captions are burned in from the word-level transcription.
  5. The final cut is stitched into an mp4 for the platform.

Editing the video after that is editing text: changed the price in a line — re-voice one phrase and regenerate its seconds.

What it costs

On current pricing: voice clone — $0.06 per 1,000 characters of voiceover, voice design — $0.05 per voice, video with a character from a photo — from $0.60 per minute on Standard 640p and from $0.80 on Lite 480p, up to $1.80 per minute on Pro. Billing is per second, no subscription: a 20-second creative costs cents, a three-minute VSL a few dollars.

For comparison: a voice actor charges from $20–50 per minute of recording, and any script change is a new session.

The legal side

You can only clone a voice you have the right to use: your own, an employee's or an actor's under contract, a client's with written consent. The same goes for the face in the photo. Synthetic content with someone else's face or voice without consent is prohibited by the terms of service and by the laws of most countries, and ad platforms ban accounts for it. The safe alternative: a designed voice and a licensed model's face.

Typical problems and how to fix them

Symptom Cause What to do
The clone "reads," no intonation The sample was read from a page Record a conversational sample
Wrong stress in a word A rare word or a name Put a stress mark in the text
Swallowed end of a phrase The line has no period at the end End the line with a period, regenerate
The mouth "sticks" shut The model doesn't suit a medium shot Switch the shot to another tier
The face "swims" in a close-up Too little detail in the source photo A sharp portrait with even light, the Pro tier

Bottom line

Voice cloning and lip sync aren't "a replacement for the actor" — they're a different way of producing: the video becomes text you can edit, translate and scale. Quality depends on three things — a clean sample, a sharp photo and short lines.

Try a voice clone and a living character in a minute: the sign-up bonus covers a test video.

Build your VSL in minutes

Get $3 for your first generations: voice clone, a lifelike presenter from one photo and captions in one cycle — pay only for seconds of finished video.

Create your first VSL
Create your first creative