The music sounds lovely. You can't remember what the narrator said. That's a sound design problem, not a voice problem.
Sound design for narration means arranging the sounds around a voice so the listener understands the story with less effort. A bed can establish a mood. A stinger can mark a turn. Room tone can keep an edit from sounding like a broken connection. Silence can make the next sentence matter.
In Vois, the built-in music and sound effects library gives you material to work with, and the multi-track timeline gives those sounds their own space. Neither is a reason to fill every gap. Start with the words and ask what the listener actually needs around them.
What should sound design do around a voice track?
Give each sound a job before you add it. Useful jobs include locating a scene, supporting an emotional shift and making a structural boundary audible. “It felt empty” is worth questioning. Empty can mean focused.
Take a narrated product walkthrough. A restrained opening cue might signal the start, while a dry voice explains an unfamiliar setting. Music can return after the explanation, when the listener no longer has to track a sequence of actions. That arrangement follows the task rather than decorating the whole recording.
Listen to the narration without accompaniment first. If the phrasing is unclear, fix the script or read. If the transition makes no sense without a dramatic effect, add the missing sentence. Sound can reinforce meaning; it shouldn't be asked to manufacture meaning that isn't there.
Our multi-track timeline guide explains the arrangement workflow. Here, the question is what deserves a track at all.
How do you choose a music bed for narration?
A music bed is a continuing layer beneath speech. Choose it while the narrator is talking, not during a solo preview. A track that sounds exciting alone might compete with every sentence once you put the layers together.
Listen for busy melodies, sharp percussion and sudden changes in energy. These aren't inherently bad, but they ask for attention. Under a dense explanation, a simpler arrangement may leave more room for the words. Under an emotional scene, movement may be exactly the point.
Browse the built-in library with a scene in mind: quiet reassurance, measured curiosity or a clear finish. Avoid treating “upbeat” as a complete brief. An upbeat sales introduction and a hopeful personal story don't necessarily need the same accompaniment.
Check the end as carefully as the opening. If the music resolves halfway through an unfinished thought, choose another passage or another bed. Let the narration decide where the scene lands. Don't rewrite a useful explanation just to obey the shape of a convenient track.
How loud should a music bed sit under narration?
Quiet enough that understanding the voice takes no extra effort. There isn't a universal fader position: different recordings arrive at different levels, and some music competes with speech even when it seems quiet.
For a specific accessibility reference, the W3C WCAG 2.2 Recommendation dated December 12, 2024 includes a Level AAA criterion for qualifying prerecorded audio-only content containing primarily speech. Its options include no background sound, allowing that sound to be turned off, or keeping it at least 20 decibels below foreground speech, subject to the criterion's exceptions. That is a scoped accessibility requirement, not a universal entertainment mixing target.
In practice, establish a comfortable voice level, bring the bed up from silence, then back it down when it starts drawing your attention. Check the quietest spoken passage, not just the energetic opening. If the listener misses a name or qualifier, the balance isn't working.
Try ordinary speakers and low-volume playback. Also compare the mix with the bed removed. If the voice suddenly feels much easier to understand, reconsider the music choice as well as its level. Turning everything up at export won't repair competition inside the mix.
Where does a stinger help a transition?
A stinger is a brief cue that tells the ear something has changed. It works best where the script already contains a meaningful turn: an introduction ends, a new topic begins or a scene closes.
Place the cue after the outgoing thought has landed and before the next essential phrase needs attention. Listen to its decay, the sound that remains after the initial hit. A small chime can have a long tail that covers the start of the next sentence.
Don't give every paragraph its own entrance music. Repeated cues can make a calm explanation feel like a sequence of interruptions. Save them for boundaries that listeners benefit from noticing.
For recurring work, choose a recognizable cue and use it for the same purpose. If it announces a chapter in one episode and a warning in another, it stops being useful shorthand. The practical test is simple: can you explain why the cue belongs at this exact point?
When should you use room tone instead of silence?
Room tone is the quiet ambient sound of a recording space without intentional speech. In a recorded voice session, it can bridge edits so the background doesn't jump abruptly from a room to absolute quiet and back again.
Use matching material where you have permission to use the recording. An unrelated ambient loop may make the edit more noticeable, not less. Listen for changes in texture at the boundaries: a fan disappearing, a distant hum changing pitch or a room suddenly feeling smaller.
Generated narration is a different starting point. A clean voice track doesn't automatically need a layer of room noise to sound credible. Add ambience when the scene calls for a setting, not because every pause supposedly requires filling.
For example, a story set outdoors may benefit from environmental sound. A software instruction may benefit from nothing beneath it. These are editorial choices. Keep scene ambience distinct from the background noise used to repair continuity in a recording.
Why is silence so useful in narration?
Silence gives an idea somewhere to land. It can separate instructions, hold a question open or let an emotional sentence finish without music explaining how the listener should feel.
Try removing the bed before an important line, rather than making the accompaniment bigger. The contrast may be enough. But check that the change sounds intentional: an unexplained cut in a recording can resemble a playback fault.
Use Vois pause nodes to shape breaks in the narration. Then listen with the surrounding tracks active. A pause in the voice isn't silence if an effect is still ringing or music is carrying the scene forward. Decide whether you want a speech break, an ambient moment or genuine quiet.
The pause nodes guide is useful when a break needs deliberate placement. Avoid prescribing the same duration everywhere. A listener processing an unfamiliar term needs a different beat from someone waiting for a punchline.
How do you assemble and check the finished mix?
Keep speech, music and effects on separate tracks while arranging. That lets you adjust a bed without regenerating the narrator, and replace an effect without disturbing an approved read. Name the tracks by purpose so the project still makes sense when you reopen it.
Work through the timeline listening for entrances, exits and overlaps. Check that no sound masks the opening consonant of a sentence, and that the ending doesn't cut off a word or musical tail. Then play the whole piece without stopping. Local fixes can feel right alone and still create a restless overall rhythm.
Use the audio export options after the balance works. Mastering can help with the final presentation, but it cannot decide which sound deserves the listener's attention. Keep a voice-only version where the delivery brief requires one, and check the exported file rather than relying only on timeline playback.
For your next voiceover project, try a smaller arrangement: a purposeful opening, space for the explanation and a clear ending. Add more only when you can name the benefit.
Make room for the words.
The Vois Team