Video Editing

Audio in Video: Mixing Dialogue, Music and Effects

Sound decides whether viewers keep watching. Learn mix priorities, dialogue levels and loudness, cleaning up noise, ducking music, and licensing it properly.

Thebes International teamPublished 8 min read

Audio in Video: Mixing Dialogue, Music and Effects

Audio mixing for video is the stage where dialogue becomes clear, music supports rather than competes, and sound effects are present without calling attention to themselves. Viewers will put up with average pictures, but they usually leave when speech is hard to follow or the music is too loud. This article covers mix priorities, levels and loudness, cleaning up dialogue, using room tone, ducking music under speech, and the general principles of licensing music and effects.

Dialogue first: setting priorities

In any speech-led video, such as ads, interviews or tutorials, the order of importance is fixed: dialogue first, then the sound effects that support the picture, then music. Every mixing decision is judged by one question: is the speech understandable the first time?

Organize your tracks before touching any settings. It makes the work easier and lets you process each type of sound in one pass:

  • Dialogue tracks: each speaker or each microphone on its own track where possible.
  • Effects tracks: specific sounds such as a door opening or a notification, plus ambience such as street or café atmosphere.
  • Music track: one or two tracks for the music and its transitions.

Remember that half of good audio is made at the recording stage, not in the edit. A microphone close to the speaker in a quiet room saves hours of repair work; see Shooting for the Edit: Footage That Makes Editing Easier.

Audio levels: what the meter numbers mean

Editing software measures level in decibels relative to full scale (dBFS). Zero is the highest possible level and everything below it is negative. Push audio past zero and you get clipping, a harsh digital distortion that can't be fully repaired later.

A common working guideline is to keep dialogue peaks at roughly −12 to −6 dBFS. That leaves safe headroom above the speech for louder moments while keeping the dialogue clearly present. The steps:

  1. Even out clip levels first: raise or lower the gain of each dialogue clip so they sit close together before adding any processing.
  2. Compress gently: a compressor narrows the gap between loud and quiet words so speech feels steadier. A ratio between 2:1 and 4:1 is a common starting point for dialogue; too much compression makes the voice flat and tiring.
  3. Put a limiter on the master: it stops any peak from exceeding the ceiling and is usually set close to −1 dBFS.

Loudness and LUFS

A peak meter tells you the highest point, but not how loud the audio feels to the ear. That is what loudness metering in LUFS is for: it measures perceived loudness averaged over the whole program.

It matters because streaming and social platforms apply loudness normalization: they turn quiet videos up and loud ones down so everything plays at a similar level. Many streaming platforms normalize to around −14 LUFS, but targets vary between platforms and can change, so check the current help pages of the platform you publish on. Broadcast television uses lower targets, such as the EBU R 128 standard used in Europe and many other countries at −23 LUFS, and each channel has its own delivery specifications, so get them in writing before you mix.

The practical upshot: a video far quieter than the target will sound weak next to others, and if you crush it with compression to make it loud, the platform will turn it down and leave you with flat, lifeless sound for nothing. Use your software's loudness meter, measure the full video, and keep true peak below about −1 dBTP.

Cleaning up dialogue

Work in this order, because each step makes the next one easier:

  1. Manual cleanup first: remove obvious unwanted sounds between sentences, such as coughs or a creaking chair, and replace them with room tone.
  2. High-pass filter: removes rumble and low vibration the human voice doesn't need. A common starting point is somewhere between 70 and 100 Hz, depending on the voice.
  3. Hum removal: electrical hum sits at the local mains frequency, 50 Hz in the UAE, Kuwait and Egypt and 60 Hz in Saudi Arabia, and its multiples. Most editors include a dedicated tool for it.
  4. Noise reduction: the tool learns the fingerprint of steady noise, such as an air conditioner or a fan, and lowers it. Go easy: overdoing it gives a metallic, underwater sound.
  5. EQ: tame harsh frequencies and add a little presence in the upper mids if the voice needs it.
  6. De-esser: softens sharp "s" and "sh" sounds when they become piercing.

Room tone

No room is completely silent; every space has a faint sound of its own from equipment, ventilation or the street outside. When you cut a sentence from an interview and leave pure digital silence in its place, viewers sense a jarring dropout even if they can't name it. The fix is room tone: roughly 30 to 60 seconds recorded on location with no talking or movement, which the editor uses to:

  • Fill gaps between sentences after cuts.
  • Smooth the join between two takes recorded at different times.
  • Give the noise reduction tool a clean noise profile.

If nobody recorded room tone, look for pauses between sentences in the footage, then copy and loop them carefully.

Music: choosing, ducking and licensing

Choosing

Choose music for the pace and mood the video needs, not only your personal taste. Songs with vocals compete directly with dialogue, so instrumental tracks or quiet sections work best under speech. Note where the track changes, such as a strong opening or a climax, and line those moments up with key points in the video. For some Ramadan campaigns and content tied to religious occasions, many brands prefer sound design, percussion or vocal-only pieces instead of music. Agree that with the client early, not on delivery day.

Ducking music under speech

When speech starts the music dips, and when it stops the music rises a little. You can do this by hand with volume keyframes for the most precise result, or use the automatic ducking tools many editors offer and then review what they did. The test is simple: close your eyes and listen. If you don't catch every word the first time, the music is still too loud. You can also cut a little of the midrange from the music with EQ to make room for the voice without lowering the whole track.

Licensing

Copyrighted music doesn't become usable because you bought it to listen to, or because you only used a few seconds. The general principles:

  • Royalty-free doesn't mean free: it usually means you pay once or by subscription with no per-use fees, under the terms of the license.
  • Commercial use has its own terms: ads and sponsored videos need a license that explicitly covers commercial use. On some platforms, business accounts only get access to a limited commercial music library, so check what your account type allows.
  • Automated matching systems: platforms such as YouTube can detect copyrighted music and restrict the video or redirect its revenue, even for short uses.
  • Keep records: store the license file or receipt with the project files for every track and sound effect.

These are general principles, not legal advice. For the wider picture on licensing, see Image Rights and Licensing; the same logic applies to music.

Sound effects

Sound effects make the picture feel physical: a cup set down on a table, a car door closing, a phone notification. They fall into specific sounds tied to on-screen events, ambience that fills a space with life, and Foley, movement sounds recorded to picture such as footsteps and rustling clothes. Keep them lower than you expect; a good effect is felt rather than noticed. Sounds for transitions and animated elements in motion graphics follow their own rules, covered in Sound Design for Motion Graphics.

Monitoring and delivery

  • Listen on at least three devices: good headphones for noise detail, desktop speakers for overall balance, and a phone speaker, because that is what most social audiences use.
  • Check in mono: some devices fold the two channels into one, and parts of a poorly balanced mix can disappear. Keep dialogue centered.
  • Sample rate: the common standard for video is 48 kHz, so set the project to it from the start.
  • Delivery: export the final mix, and keep separate dialogue, music and effects stems if the client may need other language versions or later changes. Audio codec settings at export are covered in Video Export Settings: Formats, Resolution and Bitrate.

If you also produce audio-only content, most of these techniques carry over, as covered in Podcasting: From Idea to First Episode. To see where this stage sits in the full workflow, read Video Editing: The Complete Beginner's Guide.

Audio checklist before delivery

  • Is the dialogue clear the first time on a phone speaker?
  • Do dialogue peaks sit around −12 to −6 dBFS, with no clipping anywhere?
  • Have you measured the loudness of the full video and compared it with the platform's or channel's target?
  • Are edits free of dropouts, with room tone filling gaps where needed?
  • Is the noise down without the voice turning metallic?
  • Does the music dip under speech and end deliberately rather than cutting off?
  • Is there a stored license covering this use for every track and sound effect?
  • Have you listened to the whole video in mono and on three different devices?

Related articles