• Three common levels: condition on a short reference clip, adapt the model to one speaker, or continue from recent speech context.
  • Reference conditioning is convenient; adaptation usually improves fidelity but needs more clean data and a separate model or adapter.
  • Evaluate intelligibility and speaker similarity separately. A sample can sound natural while sounding like the wrong person.

Connections