Built from a speech tokenizer, a language model over those tokens, and a token-to-speech synthesizer. Connections related_to Speech-to-Speech