VSL·TOOLS

BLOG Blog

Replace Speaker in Video & Translate with AI: Localize Video Ads

·6 min read

Every funnel has a shelf life: the creative burns out, the offer changes its price, a geo closes. This used to be where the video got reshot. Now a finished video can be edited like text: replace the speaker in the video, translate it into another language, rewrite a single line — and regenerate only the seconds that actually changed.

Let's see how it works using the speaker replacement tool in VSL·TOOLS as the example: what it can do and where the limits are.

What speaker replacement can do

Input: any finished video — an ad, a review, a testimonial, a VSL. No source files, editing project or separate tracks needed. The service takes the whole video apart:

  • finds the characters on screen — who speaks when, where the face is in close-up and where it's b-roll;
  • transcribes the speech with word-level timing — an editable transcript;
  • splits the speech into lines by speaker.

Analysis is free. From there you work with text and characters, not a timeline.

Three scenarios

  1. Replace the character. Pick a speaker on screen and put a new one in their place: one photo and 10 seconds of voice give a new face and a new voice — a face swap video for ads, except the edit, pacing, shots and b-roll stay original.
  2. AI video translation. The lines are translated, voiced with a clone of the original speaker's voice (or a new one), and the lips are re-synced to the new language — AI dubbing without a dubbing studio. Around 200 translation languages are supported.
  3. Edit a line. Rewrite a phrase as text — a new product name, price, call to action. The service re-voices it and regenerates the lips in that line only.

The scenarios combine: you can replace the speaker and translate the video at the same time.

How it works under the hood

The main principle: only what changed gets regenerated. Everything untouched goes into the final cut from the original, native audio included.

  • A line you rewrote or translated becomes a new fragment: voiceover by the clone, character lip sync, captions.
  • Lines you left alone come from the source without re-encoding or quality loss.
  • Pauses, b-roll and scenes without the replaced speaker's face — also from the original.
  • Where the old speaker's face flashes in a pause, the shot is cut: nothing to show there, and even less to hear (the old voice is still fading out in that pause).

That's why the reworked video keeps the pacing and the edit of the original: the viewer sees a different person or hears a different language, but the rhythm doesn't change.

What happens to a face on b-roll

If another replaced speaker appears in the frame (a reverse shot on the listener, say), that segment can't come from the original — it's the old face. The service puts in a shot with the new character of that specific speaker, not "whichever is nearest."

Step by step

  1. Upload the video — an mp4 up to several gigabytes, no source files.
  2. Wait for the analysis — a few minutes: characters, transcript, lines.
  3. Decide what to change. Each speaker has two options: "keep as-is" or "replace with a character." A character is a photo plus a voice clone or a designed voice.
  4. Edit the text. Translation into the language you need with one button, line edits right in the transcript. You can add lines: a new phrase goes into the video as its own scene.
  5. Voice it — line by line, any phrase regenerated on its own.
  6. Generate the lips — only for the changed lines; the blocks are assembled automatically.
  7. Assemble the final cut — the service splices the new pieces with the original; the output is an mp4, with captions if you need them.

What it costs

You pay for the seconds actually redone: rewrote one 8-second line in a three-minute video — you pay for 8 seconds of video and the voiceover characters. A full translation with speaker replacement is voiceover for the whole text plus lip sync only for the frames where the speaker is on screen, not the entire running time. Current prices are on the pricing page; analysis of an uploaded video is free.

Versus a reshoot: localizing a three-minute video for a new geo costs single-digit dollars and takes minutes — instead of a new day on set with an actor.

Where the limits are

  • Rights. You can replace a speaker and clone a voice only where you hold the rights to the source video and have the consent of the people whose faces and voices are used. A competitor's creative downloaded from a spy tool is not your material.
  • Final length. Cutting the pauses that show the old speaker's face can make the video slightly shorter than the original — deliberately: a clean assembly matters more than matching the length to the second.
  • Very small faces. If a face takes up a few dozen pixels, lip sync on it is pointless — leave such shots as b-roll.
  • Music on top. Speech is separated from the background, but the louder the music in the original, the harder a clean translation; better to have a version without music.

Typical uses in performance marketing

  • Localize video ads for new geos. One winning video → five languages with the same voice and face.
  • Change the offer. Same video, different product and price: two lines rewritten.
  • Swap the spokesperson. The actor left or burned out — a new face, the same edit.
  • A/B test the spokesperson. One video, three characters — and traffic tells you who gets trusted more.

Bottom line

Speaker replacement and translation turn a finished video into an editable document: change the text, the face or the language — only the changed seconds are regenerated, the rest stays original. It's the cheapest way to extend a funnel's life and open a new geo without a reshoot.

Upload a video for analysis for free — you only pay for the seconds you decide to redo.

Build your VSL in minutes

Get $3 for your first generations: voice clone, a lifelike presenter from one photo and captions in one cycle — pay only for seconds of finished video.

Create your first VSL
Create your first creative