AUD · Audio tools

Audio Formant Shifter

Shareable link

Settings are written to the URL as you change them. Nothing differs from the defaults yet.

Change the size a voice appears to be without changing the note it is singing. The resonances move; the partials do not.

Pitch and formants are separate

A sung note puts energy at a fundamental and at whole multiples of it. Where those partials sit is the pitch.

How loud each one is relative to the others is set by the resonances of the throat and mouth, and those sit at frequencies that barely move as the singer changes note. They are what identifies a vowel, and what tells a listener whether they are hearing a large person or a small one.

Every pitch shifter moves both at once, because scaling the spectrum scales everything in it. That is the chipmunk effect: the note went up, and the apparent size of the speaker went up with it.

Semitones and fine

Semitones runs from -12 to +12, with the fine control adding up to 50 cents either way.

Negative shifts stretch the envelope down the spectrum, which reads as a longer vocal tract and a larger speaker. Positive shifts compress it upward for a smaller one.

Small amounts are more convincing than large ones. Two or three semitones changes the character while leaving the voice believable. Past about eight it stops sounding like a person and starts sounding like a process, which is sometimes the point.

Correcting a pitch shift

Pairing this with the pitch shifter is the reason most people want it.

Shift the pitch up 5 semitones and the voice is 5 semitones high and noticeably small. Shift the formants back down 5 semitones and the resonances return to where they were, so the result is a higher note sung by a person of the original size. The same pairing in reverse gives a lower note without the growl.

How the envelope is found

The step separates the spectrum into a slowly varying envelope and the harmonic ripple underneath it, using the cepstrum: the logarithm of the magnitude spectrum, transformed again, splits the two apart because they vary at different rates across frequency.

Everything varying faster than roughly 800 Hz across the spectrum is treated as ripple and discarded, leaving the envelope. The envelope is then resampled by the shift ratio, divided back in, and the original phases are put back untouched. Nothing in that path can move a partial, which is what guarantees the pitch survives.

The consequence is a floor: a voice with a fundamental much above 800 Hz has harmonic spacing wide enough to be mistaken for envelope, and the separation degrades.

Example: a narration that needs more authority

Semitones -3, fine 0, mix 100%.

The resonances drop three semitones and the read sounds like it came from a larger person, at exactly the pitch it was recorded at, with the timing untouched.

Frequently Asked Questions

A resonance of the vocal tract. The throat and mouth amplify certain frequency regions regardless of the note being sung, and those regions are what tell a listener the size of the person speaking and which vowel they are making. They stay in roughly the same place as the pitch moves, which is why they can be shifted separately.

A pitch shifter moves the partials and drags the resonances along with them, which is why a voice pitched up sounds like a chipmunk. This moves the resonances and leaves every partial exactly where it was, so the note is unchanged and only the character of the voice moves.

Shift down for larger and up for smaller. Negative values stretch the apparent vocal tract, which reads as a bigger speaker; positive values shorten it. Three to five semitones is a noticeable change without sounding processed.

Pitch it up with the pitch shifter, then shift the formants down by the same number of semitones. The pitch shifter moved the resonances up along with the note, and shifting them back puts them where a real voice of that size would have them.

It works on anything with a clear resonant structure, so brass, reeds and bowed strings respond. It has little to do on percussion or on noise, because there is no separable envelope to move.

The step works on a spectrum rather than on samples, and it cannot analyse a window until the whole window has arrived. That is about 46 ms, and it is declared to the engine, so the preview compares against a matching delay and the export removes it from the front of the file.

Explore Our Tools

Browse all tools