AUD · Audio tools

Audio Vocoder

Shareable link

Settings are written to the URL as you change them. Nothing differs from the defaults yet.

Measure how loud a signal is in each of a set of frequency bands, then apply those levels to a synthesized tone. The result carries the rhythm and the vowels of the original with the pitch and timbre of the carrier.

How a vocoder works

The signal is split into parallel band-pass filters. Each one feeds a follower that tracks the level in that band and nothing else.

The same filter bank is applied to a carrier, and each of the carrier’s bands is scaled by the level the matching band of the source is at. Summing them back together gives a signal shaped like the source and made of the carrier.

Each carrier band is held at a steady level before the source’s envelope is applied, so the output tracks the source rather than the carrier’s own spectrum. Without that, the result would depend on how much energy a sawtooth happens to have near each centre frequency instead of on what the source is doing.

Bands

Eight, sixteen, or twenty-four bands, spread logarithmically from 120 Hz to 7.5 kHz so each covers the same musical interval.

The count is the resolution of the analysis. Fewer bands means less of the source’s spectral detail survives, which is what makes low counts sound more artificial. The classic hardware sound is at the low end of this range.

Carrier

Sawtooth has energy at every harmonic, which gives every band something to work with. This is the sung robot voice.

Pulse is hollower and reads as more nasal, with a thinner low end.

Noise has no pitch. Applying speech envelopes to noise produces an unvoiced whisper, useful when the point is anonymity rather than melody.

Carrier pitch runs from 50 Hz to 400 Hz. Set it near the speaker’s own pitch to stay close to the original, or well away from it for a deliberately mechanical result. Both waveforms are band-limited at the point where they wrap, so their harmonics do not fold back down into the bands about to be measured.

Band width and response

Band width sets the Q of every filter, from 2 to 12. Low values overlap heavily and blur the bands together; high values isolate them, which sharpens the effect and can sound comb-like.

Response is how fast each follower tracks its band, from 1 ms to 200 ms. Short values follow consonants and transients. Long values hold each level across syllables, which is smoother and less intelligible.

Intelligibility

Vocoded speech is always harder to follow than the original, and the reason is structural.

Consonants such as s, t and f are noise above 7.5 kHz, and they carry a large share of what distinguishes one word from another. A band-pass bank that stops below them cannot pass them on. Vowels, which are exactly the formant structure the bank does resolve, come through clearly.

Raising the band count, shortening the response and blending a small amount of the original back with the mix control are the three moves that recover clarity.

Example: a spoken line as a robot

Bands 16, carrier sawtooth at 90 Hz, response 12 ms, band width 6, mix 100%.

The vowels track the speaker and the whole line sits on one pitch, which is what makes it read as machine rather than as a processed person.

Frequently Asked Questions

It is synthesized inside the step. A vocoder normally needs two signals, one to analyse and one to excite, and a step in a chain only ever sees one, so the carrier is generated here from the pitch and waveform you choose.

How much detail of the original survives. Eight bands is coarse and unmistakably robotic; the words are there but the vowels blur together. Twenty-four resolves individual formants, so the speech is clearer and the effect is less extreme. Sixteen sits between the two.

The analysis only covers 120 Hz to 7.5 kHz. Anything in the source outside that range has no band to drive, so it does not reach the output at all. Follow the step with a gain stage when you need to match levels.

Sawtooth for the classic sung robot, because it has energy at every harmonic for the bands to find. Pulse is thinner and more nasal. Noise has no pitch at all, which turns speech into a whisper rather than a robot.

It sets how fast each band's envelope follower tracks the source. Short values from 1 ms to 10 ms keep consonants sharp and can sound gritty. Longer values smear the words together and make the result more musical and less intelligible.

Consonants live mostly above 7.5 kHz and in noise rather than in pitch, so a vocoder loses them. Raising the band count and shortening the response recovers some clarity. Mixing a little of the original back in with the mix control recovers more.

Explore Our Tools

Browse all tools