Reduce the room on a recording
A recording made in a live room carries the room with it. The words arrive first, and then the walls send them back, overlapping whatever is said next. Reverb remover reduces that tail while leaving the direct sound.
Three controls decide the result: how long the room rings, how much of the estimated tail is subtracted, and how far the attenuation is allowed to go.
What separates reverb from the signal
Level does not separate them. Reverb is often as loud as the voice, so a gate set to cut it cuts the words too.
Timing does. A reverberant frequency band is loud now because it was loud a moment ago. A dry one is loud because it is loud now. So each band keeps a running estimate of what its own past should still be spilling into the present, decayed at the rate the room decays, and that estimate is subtracted from what actually arrives.
Direct sound arrives before its own tail exists, so it survives. The tail is predicted by the frames that made it, so it comes off.
Room decay
This is the RT60 of the space, the time a sound takes to fall by 60 dB. It sets how fast the tail estimate decays between frames, so it needs to match the room rather than the effect you want:
200to350 ms: a treated room, a small carpeted office, a booth.400to700 ms: an ordinary room with hard floors, a bedroom, most home recordings.1000 msand above: a stairwell, a tiled bathroom, a hall, a church.
Too short and the estimate collapses before the real tail does, leaving the late reverb behind. Too long and it keeps subtracting after the room has stopped, which eats the sustain of held notes and makes them pump.
Amount and max reduction
Amount scales the estimate before it is subtracted. Around 60 to 80% takes a room down noticeably while leaving the recording intact. Above 85% the subtraction starts removing signal that overlaps the tail, which sounds hollow and watery.
Max reduction is a floor under the attenuation. No band is ever pushed down further than this, so a frame the estimate has badly overshot cannot be silenced. Lower values are safer and less effective: 8 to 12 dB is a conservative setting, 20 dB and beyond is for a genuinely bad room.
Artefacts and how to avoid them
The subtraction happens per frequency band, and a band whose gain jumps around between frames produces the fluttering, metallic sound known as musical noise. Two things hold it down here. Neighbouring bands are averaged before the gain is applied, because a real tail is broad and a single band dipping on its own is the estimate flickering rather than the room decaying. The gain is then smoothed over time, faster on the way down than on the way up.
If artefacts still appear, lower amount before raising max reduction. Amount decides how much signal the subtraction is allowed to reach; max reduction only decides how far it can go once it gets there.
In a chain
Run this before the denoiser rather than after. The denoiser learns a steady noise floor, and a reverberant tail is not steady, so it reads as signal and skews the estimate. Removing the room first gives the denoiser the stationary floor it is built to find.
An EQ afterwards is usually worth it. Reducing reverb removes low mid energy along with it, and a small boost between 150 and 400 Hz puts the weight back.
Latency
The step works on a spectrum, so it cannot analyse a window until the whole window has arrived, which is about 46 ms. That figure is declared to the engine: the preview compares against a matching delay, and the export removes it from the front of the file.