The shot changes. Your narrator is still describing the last screen. Timing narration to picture is usually a script problem before it is an audio problem: the sentence contains more information than the image has room to hold.
You don't need a recording studio to fix that. You need a scene map, a repeatable voice, and permission to cut words. Keep the picture in your video editor and build the narration alongside it. Vois handles the voice production; the video editor remains the place where you judge whether sound and picture agree.
How do you write narration to a fixed duration?
Start with the usable speaking window, not the total shot length. A scene may open on a title that needs quiet reading time, then end with a sound effect you want people to hear. Neither moment belongs in the narration budget.
Make a scene sheet with the scene ID, picture description, narration start, narration deadline, and essential message. Use the same time display as the edit. If your editor is working in frames, don't silently convert a boundary into rounded seconds and expect exact alignment later.
Write the shortest sentence that carries the message. Generate a rough read with the intended voice and measure that read. A generic words-per-minute estimate can't tell you how your selected voice will handle a product name, a list, or a deliberate pause.
Here's a worked example, not a pacing standard: a scene offers eight seconds for speech, but your first read takes ten. Don't try to reclaim the difference by speeding everything up. Remove the setup phrase, generate again, and listen with the shot. The actual recording is your ruler.
What should a scene-by-scene narration sheet contain?
Give every scene a stable identifier before you generate anything. Something like demo_settings_save survives an editorial reorder better than scene_after_intro. Keep the identifier separate from the words so a rewritten sentence doesn't acquire a new identity.
For each row, record the approved text, voice choice, pronunciation notes, target window, and status. Add a beat note such as “say save when the cursor clicks” rather than “match video.” That tells you what success sounds like.
Distinguish locked boundaries from flexible ones. A legal end card may be immovable. A screen recording might allow a longer hold. Ask the editor which is which before spending an afternoon squeezing an explanation into a space that could simply be extended.
The voiceover workflow works best when the voice producer and editor share this sheet. When you're both people, it still earns its keep: tomorrow's you won't remember which take matched the earlier cut.
Why generate narration per scene instead of per video?
Generate coherent scene-sized passages so a late picture change affects a small, reviewable piece. If the pricing screen disappears, you can remove its narration without rebuilding the introduction and conclusion.
Don't split every word into a separate clip. A voice needs enough context to form a natural phrase. Keep a thought together, especially when a sentence carries a contrast or lands on a concluding word. A scene boundary is useful only when it also makes sense to the ear.
Use consistent voice and generation choices across the project. Compare adjacent scenes before declaring either finished. One may sound much more excited than the other even when each sounds convincing alone.
Save the accepted take before experimenting. Label the replacement as a candidate until you've heard the transition into and out of it. Our multi-track timeline guide covers arranging audio; picture sync still needs a check in the actual video edit.
How do pause nodes help a line hit a cut?
A pause node creates intentional silence between spoken sections. Use it where the viewer needs to inspect something, where a reveal should land, or where the next thought belongs after a picture change.
Generate the surrounding speech first. Measure where the first phrase ends and where the next one should begin. The remaining gap is the pause you need to audition. Punctuation can shape delivery, but it isn't a dependable substitute for an explicitly placed silence.
For example, “Open Settings” might finish before the panel appears. Place a pause after that instruction, then begin “Choose your output” when the next control is visible. If the first phrase itself runs past the panel opening, adding silence won't solve anything. Rewrite or reposition the phrase.
The pause node guide explains the audio side in more detail. After any regeneration, check the result again: a fixed pause does not guarantee identical durations in the speech surrounding it.
How do you match a spoken word to a visual beat?
Choose the anchor word, then build the sentence around it. For a reveal, that might be the object's name. For a tutorial, it might be the action the viewer sees. Not every noun deserves synchronization; pick the moments where a mismatch would confuse someone.
Listen for the word's actual onset rather than lining up the beginning of the clip. Leading silence and introductory words can put the meaningful beat well after the file starts. Position the audio against the picture, play through the transition, and judge it at normal speed.
Try “Your report is ready” instead of “And after that process has finished, your report is ready.” The shorter lead-in makes the anchor easier to place and gives the image room to communicate.
Leave room after a reveal as well. If the sentence immediately rushes into the next instruction, the viewer may hear the cue without having time to understand the screen. A clean ending can do more work than an extra explanatory clause.
What should you do when the narration runs long?
Cut repeated information before changing delivery speed. Remove phrases that announce what the sentence is about to do, such as “What we're going to do here is.” Replace a noun-heavy construction with an action. “Make a selection of the destination” becomes “Choose the destination.”
Check whether the picture already supplies the missing context. If the screen clearly says Export, the narration may only need to explain which option matters. Don't narrate every visible label unless the audience needs that support.
If every remaining word is necessary, discuss a longer shot or an additional scene. A hurried warning is not an acceptable solution to a locked duration. Moving essential information into an unreadable flash of on-screen text is no better.
Only then audition a modest delivery adjustment. Listen for squeezed consonants, odd emphasis, and a sudden energy change beside neighboring clips. Keep the more natural take if the faster one technically fits but makes the instruction harder to follow.
How do you check the final narration against picture?
Export a review file using Vois audio export, then place it in the video edit. Watch the whole sequence without stopping. Mark moments where your attention has to choose between listening and reading.
Run another pass focused on boundaries: clipped opening syllables, words crossing the wrong cut, abrupt silence, and music covering instructions. Check the final picture version, not an earlier reference with slightly different shot lengths.
Keep the scene sheet beside you and mark each approved row with its take and picture version. If the editor changes a hold, reopen that row instead of assuming the old approval still applies. Archive the review export separately from the approved delivery so nobody imports the wrong one.
Make the words fit the moment, not the other way around.
The Vois Team