Video summary

Fix Muddy Dialogue (Part 2): Compression & De-Ess Without Ruining Clarity

Main summary

Key takeaways

Educational

Main ideas / lessons

  • Goal of spoken-word mixing (after EQ & noise reduction): Use compression and de-essing to manage dynamics—improving clarity and polish while preserving intelligibility and avoiding audible artifacts.
  • Risk of over-processing: Compression and de-essing can make dialogue less clear by:
    • Smashing transients (attack too fast) → consonants lose definition.
    • Excessive compression (threshold too low / ratio too high) → pumping/ducking, reduced headroom over other layers, and less “oomph” to cut through.
    • Over-deessinglisping or “Daffy Duck” / lispy artifacts.
  • Processing should be as invisible as possible: Like EQ/noise reduction, the aim is improved sound without you really hearing the processor working. You want control of peaks/sibilance, not obvious effects.

Method / workflow (detailed steps)

1) Choose where compression lives in the chain

  • Prefer putting compression after:
    • EQ
    • Noise reduction (if needed)
    • De-essing (DS)
  • Compression and de-essing interact, so order matters.

2) Track compression vs bus compression

  • Track compression (often bypassed in this approach):
    • Used differently than bus compression.
    • Can be more aggressive/bitey (faster attack) if enabled.
  • Bus compression (main recommended approach):
    • Typically very gentle, controlling peaks while keeping dialogue transparent.
    • Helps achieve professional loudness consistency without constant manual fader rides.

3) Use bus compression as a “peak catcher”

  • Mindset/example starting values:
    • Threshold: around -6 dB
    • If dialogue averages around -27 dB, the compressor mainly affects super loud peaks (≈ 10 dB above average).
    • Ratio: low (gentle)
    • Makeup gain: tiny amount (as needed)
    • Knee: set so compression engages slightly earlier if desired
      • A bigger knee = smoother start, but can increase reduction.

Attack and release (important for intelligibility)

  • For dialogue, attack and release are crucial.
  • If attack is too fast:
    • It clamps down consonant transients (e.g., B/P starts; consonants lose “bite”)
    • Result: lowered clarity; consonants sound “smooshed” and less defined
  • Preferred guidance (examples mentioned):
    • Slower attack: ~100 ms
    • Slower release: ~200 ms
  • Language note:
    • Some languages rely more on transient clarity (examples: German, Spanish, and discussed broadly Mandarin).

4) (Optional) Track compression approach (if used)

  • When using track compression, the described strategy:
    • Faster attack: ~5–10 ms (more aggressive/bitey)
    • Threshold: around average levels (example: -24 dB) or left higher
    • Ratio: may be increased vs bus
    • If available, use range/max reduction limiting:
      • Keep reduction small so shouts don’t become crunchy
      • Example intent: cap compression around ~6 dB max

5) Listen critically; avoid audible pumping/gain reduction

  • The speaker contrasts “bad compression” vs “good compression” by changing threshold/ratio/attack and comparing.
  • Target:
    • When bypassed, dialogue should already feel contained
    • You should not hear pumping or obvious gain reduction
    • Compression should be transparent—you should hear the story, not the effect

6) Know when compression hurts clarity (esp. with music/SFX)

  • Mechanism described:
    • Heavy compression reduces dialogue peak headroom.
    • Music/SFX fill up the low end, so dialogue peaks can’t rise above other layers.
  • Guidance for action-heavy sequences:
    • Consider bypassing bus compression, or
    • Raise the threshold / lower the ratio so dialogue retains dynamic range and cuts through.

De-essing (DS) methodology

When to use a de-esser

  • Use a DSER when every “S” is consistently sharp and irritating.
  • If it’s only an occasional stray “S,” use volume automation / fader ducking instead.
  • Notes:
    • DS is typically done track-by-track because sibilance can be mic-dependent.
    • Mentioned mic examples: certain lavalier mics (e.g., Sennheiser / similar models), contrasted with DPAs described as smoother.

How to set up DSER (track-by-track)

  • Example using FabFilter Pro-DS:
    1. Solo the track/clip with sibilance.
    2. Loop only the problematic “S” moment (don’t process normal speech “air”).
    3. Use the DSER’s band-pass / frequency range to target where the S lives.
      • Adjust for voice gender differences (female sibilance often higher frequency, generally).
    4. Set DSER threshold so it triggers only on harsh S content (avoid triggering on air/normal consonants).
    5. Adjust range/depth:
      • Enough to remove irritation, but not so much that it causes lisping.
      • Too much reduction creates “lisp/Daffy Duck”-like artifacts.
    6. Verify in the composite mix (unsolo context):
      • The S should be less harsh while intelligibility stays natural.
    7. Keep it clip-specific when possible:
      • Bypass except for the sections/moments needing DS.

Demonstrated “wrong way” to DS

  • Making the DSER too broad/aggressive:
    • Lower threshold so it triggers on too many sounds
    • Or set maximum reduction too high
  • Result:
    • The S gets “smacked down” excessively → lispy artifacts.

Practical target numbers / starting points mentioned

  • Compression target: aim for ~3–6 dB reduction on the biggest loud peaks (not constant reduction) for transparency.
  • Example bus behavior:
    • Threshold around -6 dB
    • Dialogue averages ~-27 dB so compression mainly hits louder peaks
  • Track compression (if used):
    • Attack 5–10 ms
    • Cap reduction via range (example intent: up to ~6 dB max)
  • Bus attack/release for smoothing/leveling (example):
    • Attack ~100 ms
    • Release ~200 ms
  • DSER depth:
    • A specific case may require around ~15 dB reduction, but going that far risks making sibilance unnatural.

Closing principles

  • Do the least processing necessary to get the biggest result.
  • Train your ears:
    • Compression is less obviously identifiable than EQ; listen for cues like “squished/crushed/pumping.”
  • Avoid the “fader monkey” mindset:
    • Don’t apply changes automatically just because you can—apply them based on what the source material needs.
  • End objective:
    • Support the story with controlled dynamics:
      • not so dynamic that listeners adjust volume,
      • not so crushed that it causes fatigue or poor intelligibility,
      • not so de-essed that it turns lispy.

Speakers / sources featured

  • Speaker (primary): Unnamed instructor/mixer (referenced only by role; “Chris Nolan” mentioned as a joke/example, not as a featured speaker).
  • Software / plugins mentioned:
    • Avid Pro Tools (Pro Compressor, track/bus workflow, ProTools automation approach)
    • FabFilter Pro-C 2
    • FabFilter Pro-DS
    • MCDSP SA2 (de-esser mentioned as an alternative; details briefly described)
    • iZotope (de-essing alternatives mentioned generally)

Original video