AUD · Audio tools

Audio Pitch Correction

Shareable link

Settings are written to the URL as you change them. Nothing differs from the defaults yet.

Tune a vocal to a key

Pitch correction measures the note being sung, compares it with the scale you chose, and moves the audio onto the nearest note of that scale. The file keeps its original length. Only the pitch changes.

Four things decide the result: the key, the scale, how far each note is pulled, and how long the pull takes.

Key and scale

The scale is the set of notes a sung pitch is allowed to land on.

Chromatic allows all twelve semitones. Every note is rounded to the nearest one, which needs no knowledge of the song and suits speech, bass lines, and single sustained notes.

Major and Minor allow seven notes, and the two pentatonic scales allow five. A smaller set makes a bigger correction, because a pitch that falls between two allowed notes has further to travel. That is the point: in A minor, a G sharp sung slightly flat is corrected up to A rather than being confirmed as a G sharp that does not belong in the key.

Getting the key wrong is worse than using chromatic. A correct key fixes notes that chromatic would leave alone; a wrong key moves correct notes somewhere they should not be.

Strength and retune speed

Strength is the fraction of the measured error that gets corrected. At 100% the note lands exactly on the scale. At 50% it moves halfway, which tightens a performance while leaving its character. At 0% the step passes the audio through.

Retune speed is how long that correction takes. It is the control that decides whether the result sounds corrected or sounds processed:

  • 1 to 10 ms: the pitch snaps before the note has finished starting. Vibrato and scoops are flattened out, which produces the hard, stepped sound the effect is known for.
  • 20 to 50 ms: fast enough to catch a note as it settles, slow enough to leave the attack.
  • 60 to 200 ms: the correction arrives behind the singer. Vibrato survives, and only the sustained centre of each note is moved.

Detection limits

The detector uses a harmonic product spectrum, which multiplies the spectrum against decimated copies of itself. A fundamental is the frequency its own harmonics agree on, so this finds it even when the fundamental is quieter than the partials above it, which is normal for voice and for a small speaker.

It only tracks one pitch at a time, between 40 Hz and 1200 Hz. Two notes at once means one of them is chosen and the whole signal moves with it. When no pitch is confident enough, such as during a consonant, a breath, or a cymbal, the last correction is held rather than released, so a word does not detune halfway through.

Correction is bounded to 7 semitones. Past that the detector has almost certainly locked onto the wrong harmonic, and holding the previous amount is safer than transposing a phrase by a fifth.

In a chain

Correction works best on a signal that has already been cleaned. Put the de-esser and the denoiser before it, because sibilance and hiss both give the detector something to chase.

Formant shifting after correction is what keeps a large correction from sounding thin. Moving pitch drags the vocal tract resonances along with it, and shifting the formants back down restores the size of the original voice.

Latency

The step works on a spectrum rather than on samples, so it cannot analyse a window until the whole window has arrived, which is about 46 ms. That figure is declared to the engine: the preview compares against a matching delay, and the export removes it from the front of the file.

Frequently Asked Questions

It measures the pitch of whatever is playing about 90 times a second, works out the nearest note in the key you chose, and shifts the audio by that difference. The length of the file never changes, only the pitch.

The pitch shifter moves everything by a fixed interval you choose, so an out of tune note stays out of tune at its new position. This measures each note and moves it by whatever amount lands it on the scale, which is a different amount every time.

Chromatic snaps to the nearest semitone of any kind and needs no decision, which suits speech, bass and single held notes. Pick a specific key and a major or minor scale when you know the song, because it stops a flat note being rounded onto a semitone that is not in the key.

Retune speed is how long the correction takes to travel from the sung pitch to the target. Under about 10 ms the pitch arrives before the note does, which removes the small scoops and vibrato that make singing sound human. That hard sound is the well known effect, and longer settings around 60 to 150 ms correct the note while leaving the performance intact.

Not usefully. It looks for one fundamental at a time, so it works on a single voice or instrument. A full mix has several notes sounding at once and the detector will pick one of them, then move the whole mix by that amount. Percussion has no fundamental to find, so the step holds its last correction and leaves it alone.

Explore Our Tools

Browse all tools