Five controls over five processes. The character preset picks the pitch and the throat size that make a voice somebody else, and the sliders let you take it somewhere the preset did not go.
Pitch and timbre are separate
A voice carries two things a listener reads as size. The first is pitch, the rate the vocal folds vibrate at. The second is the set of resonances the throat and mouth impose on top, called formants, which stay roughly where they are whatever note you sing.
Most cheap voice effects move only pitch. That is why they sound like a tape played at the wrong speed: the formants move with the note, and a listener hears a record running fast rather than a different person.
Here the two are separate processes, so you can drop the pitch 5 semitones and the formants 3, or raise pitch and lower formants at once, which is the combination that reads as a large body making a high sound.
The characters
| Character | Pitch | Timbre | Colour on top |
|---|---|---|---|
| Natural | 0 | 0 | none |
| Deep | -5 st | -3 st | light drive, 2.5 dB at 2.2 kHz |
| Bright | +6 st | +4 st | none |
| Robot | 0 | 0 | 110 Hz ring modulation at 85% |
| Alien | +4 st | -7 st | 420 Hz ring modulation at 45% |
| Monster | -9 st | -6 st | heavy drive, 40 Hz ring modulation |
| Radio | 0 | 0 | band-limited 450 Hz to 3 kHz, driven |
Pitch and Timbre add to whatever the character set, rather than replacing it. Monster at Pitch +2 is still a monster, two semitones up.
Ring modulation is what makes a robot
A ring modulator multiplies the signal by a sine wave instead of mixing one in. Every frequency in the voice comes out as a pair, one at the sum with the carrier and one at the difference, and neither is harmonically related to the original. That inharmonic pair is the metallic quality; no amount of EQ produces it, and no amount of EQ removes it.
At 110 Hz the sidebands sit close enough to the original partials to read as a buzz on the voice. Push the character to Alien and the carrier moves to 420 Hz, far enough that the sum and difference tones become audible as separate pitches.
Intensity and the EQ corners
Intensity scales the whole character toward neutral, including the EQ. This matters more than it sounds: if the corners stayed at the character’s own values, Radio at 0% intensity would still be band-limited to a telephone, which is the opposite of what a neutral setting should mean. The high-pass corner pulls back to 20 Hz and the low-pass to 20 kHz as Intensity falls.
Tone is a single tilt across two of the four EQ bands. Negative values lift a low shelf at 250 Hz and cut the presence peak; positive values do the reverse. It is there to put back weight or edge that a large shift took away.
Latency and tails
Both shifters report the latency of their analysis window, and the chain adds them together. Live preview delays the bypass path by the same amount so the A/B comparison stays aligned, and the export removes the leading delay and renders the tail frames after the source ends. You do not need to compensate for anything by hand.
Where it stops
This changes an existing recording. It does not synthesize speech, and it does not clone a specific person’s voice from a sample. Both of those need a trained model.
Consent matters here. Disguising your own voice in a recording you made is one thing; passing off altered audio as somebody else is another, and in many places it is illegal. The tool has no way to know which one you are doing.