Sound designVocals

Talkbox Versus Vocoder

Both make an instrument talk. One uses your mouth as the filter, the other uses your voice as a control signal.

All articles

Same Result, Opposite Mechanism

They get filed together because they produce the same effect: an instrument that appears to speak. Underneath, they have almost nothing in common. One is acoustic and happens in your mouth. The other is analysis and happens in a filter bank. Knowing which is which is the difference between getting the sound and fighting the tool.

The Talkbox Is Acoustic

A driver takes the instrument's signal and sends it up a plastic tube into your mouth. Your mouth — tongue position, jaw, lips — filters that sound the same way it filters your own voice. A microphone in front of your face picks up the result.

Your vocal cords are not involved. You mouth the words silently; the pitch comes entirely from the keyboard. That is the part people get wrong most often, and it is why singing along makes the effect worse rather than clearer.

The Vocoder Is Analysis

Your voice goes into a bank of filters that splits it into frequency bands and measures how much energy is in each one, moment by moment. Those measurements become an envelope for each band. The same band structure is then applied to a carrier signal, usually a synth.

Your voice never reaches the output. Only its spectral shape does. This is why a vocoder sounds identical whether you whisper or shout the same phrase, and why the character of the result belongs entirely to the carrier.

The Practical Differences

  • A talkbox is a physical object with a tube and a microphone, so it captures room tone, mic character and bleed from the instrument's own driver. A vocoder is clean and repeatable.
  • Talkbox formants move the way a real mouth moves, because they are a real mouth. Vocoder intelligibility depends on how many bands the filter bank has.
  • A vocoder needs a harmonically rich carrier. Feed it a sine and almost nothing comes through — there is no energy in most bands to modulate.
  • Unvoiced consonants have no pitch, so they carry badly through a pitched carrier. Most vocoders pass a noise or high-band component from the voice to keep S and T audible.
  • In both cases your own vocal pitch is irrelevant. The melody is played, not sung.
  • A talkbox has to be recorded in real time with the part played live. A vocoder can be re-rendered after the fact with a different carrier.

Getting Either One To Speak Clearly

Give The Carrier Harmonics

A saw or a pulse wave, ideally a detuned stack, gives the filter bank something to work with across the whole spectrum. Rich carriers speak; pure ones mumble.

Over-Articulate

Say it, do not sing it, and exaggerate every mouth shape. Both techniques throw away a lot of the cues that make speech intelligible, so you have to supply more of them than feels natural.

Mind The Consonants

Blend a small amount of the dry voice, high-passed so only the sibilance and transients come through. It restores the consonants without breaking the illusion that the instrument is doing the talking.

Play Chords, Not Single Notes

Both techniques get dramatically more legible with wide, sustained chords than with a moving single-note line. Space between notes is where the phrase falls apart.

Which One For Which Sound

The talkbox reads as human because it is: the resonances moving through it are an actual vocal tract. It suits lead lines that need to feel performed, and it carries the imperfection that makes them feel that way.

The vocoder reads as machine, and it is far better at chords, at holding a phrase perfectly steady, and at being recalled six months later exactly as it was. Choose based on which of those two characteristics the track needs, not on which one is easier to set up.

Press & Media

You may also like

Our latest news

More press & Media