Watch video · 2:02
Module 04 · Script & screenplay · Lesson 3

How to Check Your Script's Pace with AI Voice (Syllables per Second)

Turn your script table into a voice prompt, generate the audio, listen and note where it rushes or drags, fix the speed, then re-time the script.

The script reads fine on paper. Then the video comes back and the first line flies past, or the middle drags. By then you've paid for the video.

Pace is easy to fix in audio and expensive to fix in video. So get it right by ear first. It takes five steps:

Tap a step to jump to it.
Why you shouldn't skip this step

1. Turn your script table into a voice prompt

Start from the table you made in the script lesson. Paste it into Claude or ChatGPT with the prompt below, and you get a ready-to-paste prompt for your voice tool, with each line's emotion and pauses carried across.

Voice toolHow the emotion goes inWhere to use it
Gemini 3.8 Flash text-to-speechNotes under DIRECTOR'S NOTES, only the lines under SCRIPTGoogle AI Studio (free) or the Gemini text-to-speech app on OpenDirector
ElevenLabs v3A tag before each line, like [warm] or [excited], and "..." for a pauseElevenLabs, with a voice you pick from its library
Keep the Gemini layout exactly as it is: without the DIRECTOR'S NOTES header, the voice reads your notes out loud.
Voice prompt · from your script table
Turn my script table into a voice prompt for [Gemini 3.8 Flash text-to-speech / ElevenLabs v3].

Script table: [paste the table from the script prompt]
Voice: [who is speaking: age, gender, where they're from]
Accent: [Indian English / American / British / Hindi]

Rules:
- Use only the Line column as spoken text. Never include times, jobs, SPS or visual cues.
- Carry each line's emotion and pace from the Speaker column, and every <Pause> mark.
- Keep numbers, symbols and acronyms written as words, exactly as in the table.

For Gemini 3.8 Flash text-to-speech, use exactly this layout:
### DIRECTOR'S NOTES (do not read these aloud)
The voice, the accent and the overall tone, then the emotion and pace line by line, then "Do not rush."
### SCRIPT (read only this)
The lines, one per line, nothing else.

For ElevenLabs v3:
- Only the lines, one per line, each starting with its emotion as a tag in square brackets, like [warm], [excited] or [whispers].
- Pauses as "..." inside the line.
- Above the prompt, one line describing the voice, so I can pick it in the voice library.

Print the prompt in one block, ready to paste.
Claude · ChatGPTGemini text-to-speech or ElevenLabsFree

2. Generate the audio

Copy the voice prompt Claude or ChatGPT gave you. Open a text-to-speech app: the Gemini text-to-speech app on OpenDirector, Google AI Studio (free) or ElevenLabs. Paste the prompt, pick the voice and click Generate.

Generate one full take. Keep the same voices for every take, so you're comparing pace, not voices.

The loop: script, voice prompt with emotion, generate, check the pace, lock and re-time.

3. Listen, and write down what you hear

Listen once, all the way through, with your eyes closed. With no pictures, you hear everything the visuals would hide: a rushed line, a jump, a sentence that sounds written rather than said. Here's the same 15-second script, voiced twice:

Squeezed into 10 seconds: 7.8 SPS
Natural, 13.5 seconds: 6.0 SPS

The natural take sits at 6.0 SPS, right for talking to camera. Squeeze the same words into a 10-second slot and it jumps to 7.8: the hook is gone before it lands.

As you listen, write down the time of every spot that feels fast or slow. Then listen once more for the pace across the whole video: does the middle speed up, does the end drag?

This matters most with two or more characters. Each one has its own voice, and voices run at different speeds, so one character can sound rushed next to the other even when both lines look fine on paper. Note that too.

0:04–0:07 · the hook feels rushed
0:12 to the end · drags a little
Character B · faster than Character A all the way through

4. Speed it up or down

Voice notes set the emotion, not the pace. Asking Gemini for a "very fast, rushed" read barely moved it (6.2 SPS against 6.0). So change the speed of the audio itself. Try it on the natural take:

45678
6.0 SPS · 13.5 s · natural for talking to camera
Pick a speed, press play. The browser keeps the pitch, so it still sounds like the same person.

Now fix it with your notes. Upload the audio to Claude, ChatGPT or OpenDirector, paste the prompt below with your notes, and it gives you back the final audio, plus a list of what it changed. You'll need that list in step 5.

Speed prompt · with your listening notes
I've attached the voice take for my ad, and my notes from listening to it.

My notes (times in the take):
[for example:
0:04–0:07 feels rushed: slow it down
0:12 to the end drags: speed it up
Character B sounds faster than Character A all the way through: slow B down a little]

Change the speed of the audio to fix each note:
- Keep the pitch, so it still sounds like the same person (use a tempo change, like ffmpeg atempo, never a plain speed-up).
- Avoid big jumps: a voice pushed much faster or slower starts to sound weird. Up to about 20–30% is usually fine. If a note needs more, tell me, and say whether a different voice or a change to the voice prompt would fix it better.
- Join the parts back without clicks or gaps, and keep the pauses between lines.

Give me the final audio file, and a list of what you changed: start and end time, and the speed, for each stretch.
Claude · ChatGPT · OpenDirectorAttach the audioFree

Prefer to do it yourself? Any tool that keeps the pitch works: the speed setting in CapCut, Audacity's Change Tempo, or ffmpeg's atempo.

Either way, avoid big jumps. Push a voice too fast or too slow and it starts to sound weird. Up to about 20–30% is usually fine. If it needs more than that, speed isn't the real fix: pick a different voice, or update the voice prompt (a calmer or livelier delivery) and generate again.

What you're aiming for:

Kind of videoTarget SPSFeels
Explainer, course, brand film4.5 to 5.5Calm and clear
UGC, testimonial, talking to cameraabout 6A real person chatting
Energetic promo or hook6.5 to 7Quick but clear
Any videoover 7.5Too fast. Cut words.
For English. Hindi runs about 10% faster; see the language table in the script lesson.

Don't flatten it to one number. These targets are for the ad as a whole. An ad where every line runs at exactly the same pace sounds flat. The weighty line should slow down, and the excited one should lift.

5. Update the script to the new timing

A new speed moves every line, and the video gets cut to those times. So update the script before you make any video. Paste your table and the list of changes from step 4 into Claude or ChatGPT:

Re-time prompt · after a speed change
Here is my script table and the final audio I've locked.

Script table: [paste it]
Speed changes: [paste the list of changes from step 4, for example "4.0–7.0s at 0.9x, 12.0s to the end at 1.1x", or one speed for the whole take]
Length of the final audio: [seconds]

Update the table to the new timing:
- For each changed stretch, divide its times by its speed (1.1x faster means divided by 1.1), then shift every later line by the time gained or lost.
- Recompute each line's SPS and the average, and check the total matches the length of the final audio.
- Flag any line now over 7.5 SPS or under 4.5 SPS, and any stretch where every line sits at exactly the same pace.
- Keep the lines, jobs and emotions exactly as they are.

Print the updated table in the same columns.
Claude · ChatGPTFree

Here's the natural take re-timed at 1.1×:

LineAt 1.0×At 1.1×
You're drinking your apple cider vinegar wrong.0.3–2.6s0.3–2.4s
Most people take it at night, when it can't do much.3.1–5.2s2.8–4.7s
Take one tablet in water before breakfast instead.5.6–8.2s5.1–7.5s
It tastes like green apple, not vinegar.8.6–10.6s7.8–9.6s
Tap the link and get your first week free.11.1–13.0s10.1–11.8s
About 6.6 SPS at 1.1×, still inside the talking-to-camera range.

Try it now

  1. Paste your script table and the first prompt into Claude or ChatGPT, and copy out the voice prompt.
  2. Generate one take in Google AI Studio (free), the Gemini text-to-speech app on OpenDirector, or ElevenLabs.
  3. Listen with your eyes closed. Write down the times that feel fast or slow, and any character who sounds more rushed than the others.
  4. Upload the audio and your notes to Claude, ChatGPT or OpenDirector with the speed prompt, and keep the final audio it gives back.
  5. Run the re-time prompt with the list of changes, and use the new times for your screenplay.

Quick answers

What is SPS?

Syllables per second while the voice is speaking. About 6 feels natural for talking to camera.

Why syllables and not words?

Words vary in length. "Unbelievable" is one word but five syllables.

Can I just speed up the AI voice?

Yes. Up to about 20–30% faster or slower, with a tool that keeps the pitch, usually still sounds natural. Beyond that it can sound weird, so pick a different voice or update the voice prompt instead.

Gemini text-to-speech or ElevenLabs?

Gemini 3.8 Flash text-to-speech is almost free and follows emotion notes well. ElevenLabs gives you a bigger choice of voices. The voice prompt above works for both.

Why write the notes under DIRECTOR'S NOTES?

That layout stops Gemini from reading your instructions out loud.

My two characters sound like different speeds. Why?

Each character has its own voice, and every voice has its own natural pace. Note which one sounds rushed and slow just that voice down in step 4.

Keep going

Get the pace right before the video.

Voice, adjust and re-time your script on OpenDirector.

Start in OpenDirector →
© 2026 OpenDirector — AI video and image studio. A SurgeGrowth product. Privacy · Terms · Cookie settings