VSL·TOOLS

BLOG Blog

How to Make a VSL Video with AI in 15 Minutes: Step-by-Step Guide

·7 min read

Not long ago, making a VSL video meant a studio, an actor and a week of editing. Now the whole cycle — from script to a finished video with a living spokesperson — fits in one service and takes minutes. Below is a step-by-step guide to how to create a VSL with AI, using the VSL·TOOLS studio as the AI VSL generator: what to prepare, how generation goes and where people usually slip up.

What you need to start

  • A script. The text of the video, broken into lines. No text yet? See VSL structure and formulas.
  • One photo of the spokesperson. Frontal or three-quarter view, the face taking up a good part of the frame, no harsh shadows or glasses with glare. The photo must be yours or of a person who has consented to their face being used.
  • 10 seconds of voice. A clean recording without music or noise: a voice message from a phone will do. No sample? A voice is designed from a description.

That's it. No green screens, no shoots, no editing software.

Step 1. Paste the script

Create a VSL project and paste the text. The service splits it into lines and works out who's speaking: a two-character dialogue is marked up automatically, a monologue gets one speaker.

Lines are the unit of all further work. Each one is voiced, edited and regenerated separately, so keep them short: one or two sentences. Long paragraphs the service splits by sentence itself.

Step 2. Create the character: an AI avatar from one photo and a voice

For each speaker — one photo and a voice.

Voice clone. Upload a 10-second sample — the service transcribes it and learns the timbre, manner and pace. The clone speaks any language: English, Spanish, German, Portuguese and dozens more — even if the sample was recorded in just one.

Voice from a description. No sample? Choose the character: gender, age, timbre, mood. The AI creates a unique voice no competitor has and that isn't tied to a real person.

Tip: if the video is for a specific geo, pick a sample from a native speaker — the clone inherits the accent of the original.

Step 3. Voice the lines

Click Voice all. Each line is generated separately — the main advantage over "one file for the whole video": don't like the intonation of the third phrase — regenerate only that one, the rest stays untouched.

What you can edit in the built-in editor:

  • Stress marks. Names, brands and rare words get a stress mark right in the text — the service keeps a database of problem words and inserts the marks itself.
  • Pauses and pace. Break a line with a period or a comma where you need a pause; emotional phrases with "!" and "?" are better kept as a separate line.
  • Volume and takes. Several takes of one phrase — and pick the best.

Voiceover is priced per character of text (pricing): a three-minute video is roughly 3,000 characters, which means cents.

Step 4. Generate the video: photo to video AI

When the voiceover is ready, the service builds the footage: the character from the photo speaks the text, the lips hit every word, and scenes change like in a real edit — wide shot, close-up, a different angle. It's more than a talking photo AI with a moving mouth: the shots change as if several cameras had filmed the take.

Pick the generation model for the task:

Tier What for What sets it apart
Lite Testing hypotheses, high-volume funnels Fast and cheap, stable picture
Standard Main videos, medium shots Natural gestures and facial expressions
Pro Close-ups, final versions Maximum facial detail

The service can pick the model itself by segment length: short lines go better to Standard, long monologues to Lite. You pay per second of finished video, so a 30-second hook test costs a few dozen cents.

Step 5. Add captions

Most of the feed is watched with the sound off, and a hook without captions simply doesn't land. Captions take two clicks: automatic word-level transcription, a style (background box, outline, active-word highlight) and burn-in. The preview shows exactly what ends up in the final cut.

Step 6. Assemble the final cut and check it

The final assembly stitches the scenes, levels the audio and outputs an mp4 for the platform. Before publishing, check three things:

  1. The first 10 seconds — does the hook read with captions and no sound?
  2. The seams between lines — any swallowed word endings? If so, re-voice the short line with a period at the end.
  3. The face in close-ups — does it "swim"? If the shot is a close-up on Lite, switch it to Pro.

What it costs

On current pricing, a 3-minute video with one spokesperson, a voice clone and captions comes to roughly this: voiceover of ~3,000 characters — about $0.20, captions — about $0.15, video — from $1.80 on Standard 640p ($2.40 on Lite 480p) to $5.40 on Pro. In total, two to six dollars for a finished VSL versus hundreds for a shoot — and no reshoot when the script changes.

The sign-up bonus covers your first test video, and analyzing a finished video for speaker replacement is free altogether.

Common mistakes

  • A bad photo. A blurry face, a small image, harsh side light — the character comes out "plastic." Use a sharp portrait with even light.
  • A noisy voice sample. Music or a second voice in the recording will end up in the clone. 10 seconds of clean speech beat a minute with noise.
  • Huge lines. An 800-character paragraph is hard to edit and regenerate; keep lines under 200 characters.
  • Stress marks on ordinary words. Marks are only for names, brands and rare words — on common words they sound like a foreign accent.
  • One video for all geos. Translate and re-voice with the same clone for each geo — that way the viewer doesn't feel the substitution.

What's next

Once the first VSL is ready, scale: 5–10 hook variants on one body of the video, a second spokesperson, a different length. A finished video can have its speaker replaced or be translated into another language, regenerating only the seconds that changed.

Create your first VSL right now — sign-up takes a minute.

Build your VSL in minutes

Get $3 for your first generations: voice clone, a lifelike presenter from one photo and captions in one cycle — pay only for seconds of finished video.

Create your first VSL
Create your first creative