- Turns a waveform into a compact stream of learned audio tokens and can decode those tokens back into sound. Unlike a text tokenizer, it must carry speaker identity, pitch and timing.
- Codecs make speech generation look more like language modelling: predict token IDs, then decode them into audio.
Connections
- part_of Speech-to-Speech
- uses Residual Vector Quantization