• Combines a text model, reference-audio tokenizer, speaker embedding and a flow-matching acoustic generator. The reference clip conditions both who is speaking and how they speak.
  • Useful example of why cloning quality depends on more than one speaker vector: reference speech can also carry rhythm and expression.

Connections