AUD · Audio tools

Audio De-esser

Sibilance is the burst of energy an s, t, or sh puts into the 5 to 9 kHz region. Close microphones and bright compression push it further ahead of the rest of the voice, and lowering it with an EQ costs the whole take its top end. A de-esser only reacts while the harsh band is actually loud.

Frequency

The frequency control sets the crossover the detector listens above. The range covers 2 to 14 kHz on a log scale.

To find it, raise the range to its maximum and sweep. The setting is right when the s sounds duck hard and vowels stay put. Then bring the range back to something you would use.

Threshold and range

Threshold is the level the sibilant band has to exceed before anything happens. Range caps how far the gain is allowed to fall once it does.

A de-esser is a limiter on that band, not a compressor with a ratio: everything above the threshold comes back to it. Range is what stops that from being absolute. At 6 dB the harshest consonants are 6 dB quieter and still audible, which is usually what you want on a lead vocal. At 20 dB they are pushed most of the way out of the take.

Release

Release decides how quickly the gain returns after the consonant passes. The attack is fixed at 1.5 ms, because a slower one would let the front of every s through before the gain moved.

Short release times of 40 to 80 ms suit fast speech. Longer settings hold the reduction across a whole word, which sounds smoother on a sustained note but dulls the syllables after it.

Split band vs wideband

In split band mode the band above the crossover is reduced and everything under it is passed through untouched, so a range of 0 dB is bit-for-bit the source file.

Wideband mode applies the same gain to the whole signal. On a solo vocal that reads as the take briefly stepping back rather than as a filter opening and closing, which some engineers prefer. On a busy mix it also ducks the instruments around the voice.

Example: a bright podcast take

Frequency 6.5 kHz, threshold -30 dB, range 8 dB, release 60 ms, split band.

That is enough to stop the s sounds from stinging in headphones while leaving the consonants intelligible on a phone speaker.

Frequently Asked Questions

It watches one high band for sibilance and pulls the level down only while that band is too loud, so the rest of the voice is left where it was. It reads WAV, MP3, M4A, AAC, Ogg, Opus, FLAC, AIFF, and WebM audio, and exports in the format the source arrived in.

Most sibilance sits between 5 and 9 kHz. Deeper voices usually land near 5 to 6 kHz and brighter ones closer to 8 kHz. Raise the range to 20 dB, sweep the frequency until the effect is obvious on the s sounds, then lower the range again.

Split band pulls down the band above the crossover and leaves everything below it untouched. Wideband ducks the whole signal when sibilance is detected, which is how most hardware de-essers work and can sound more natural on a full mix.

The range is too wide or the frequency is too low. Consonants need some energy above 5 kHz to read as an s at all. Try 4 to 6 dB of range first and raise the frequency by 1 kHz.

Yes. Select the range on the waveform, apply it to that range, and the crossfade at each edge keeps the transition from stepping. The rest of the file comes out unchanged.

Explore Our Tools

Browse all tools