• Turns a waveform into a compact stream of learned audio tokens and can decode those tokens back into sound. Unlike a text tokenizer, it must carry speaker identity, pitch and timing.
  • Codecs make speech generation look more like language modelling: predict token IDs, then decode them into audio.

Connections