- Combines a text model, reference-audio tokenizer, speaker embedding and a flow-matching acoustic generator. The reference clip conditions both who is speaking and how they speak.
- Useful example of why cloning quality depends on more than one speaker vector: reference speech can also carry rhythm and expression.
Connections
- part_of Text-to-Speech
- related_to Voice Cloning