Measure how loud a signal is in each of a set of frequency bands, then apply those levels to a synthesized tone. The result carries the rhythm and the vowels of the original with the pitch and timbre of the carrier.
How a vocoder works
The signal is split into parallel band-pass filters. Each one feeds a follower that tracks the level in that band and nothing else.
The same filter bank is applied to a carrier, and each of the carrier’s bands is scaled by the level the matching band of the source is at. Summing them back together gives a signal shaped like the source and made of the carrier.
Each carrier band is held at a steady level before the source’s envelope is applied, so the output tracks the source rather than the carrier’s own spectrum. Without that, the result would depend on how much energy a sawtooth happens to have near each centre frequency instead of on what the source is doing.
Bands
Eight, sixteen, or twenty-four bands, spread logarithmically from 120 Hz to 7.5 kHz so each covers the same musical interval.
The count is the resolution of the analysis. Fewer bands means less of the source’s spectral detail survives, which is what makes low counts sound more artificial. The classic hardware sound is at the low end of this range.
Carrier
Sawtooth has energy at every harmonic, which gives every band something to work with. This is the sung robot voice.
Pulse is hollower and reads as more nasal, with a thinner low end.
Noise has no pitch. Applying speech envelopes to noise produces an unvoiced whisper, useful when the point is anonymity rather than melody.
Carrier pitch runs from 50 Hz to 400 Hz. Set it near the speaker’s own pitch to stay close to the original, or well away from it for a deliberately mechanical result. Both waveforms are band-limited at the point where they wrap, so their harmonics do not fold back down into the bands about to be measured.
Band width and response
Band width sets the Q of every filter, from 2 to 12. Low values overlap heavily and blur the bands together; high values isolate them, which sharpens the effect and can sound comb-like.
Response is how fast each follower tracks its band, from 1 ms to 200 ms. Short values follow consonants and transients. Long values hold each level across syllables, which is smoother and less intelligible.
Intelligibility
Vocoded speech is always harder to follow than the original, and the reason is structural.
Consonants such as s, t and f are noise above 7.5 kHz, and they carry a large share of what distinguishes one word from another. A band-pass bank that stops below them cannot pass them on. Vowels, which are exactly the formant structure the bank does resolve, come through clearly.
Raising the band count, shortening the response and blending a small amount of the original back with the mix control are the three moves that recover clarity.
Example: a spoken line as a robot
Bands 16, carrier sawtooth at 90 Hz, response 12 ms, band width 6, mix 100%.
The vowels track the speaker and the whole line sits on one pitch, which is what makes it read as machine rather than as a processed person.