- Three common levels: condition on a short reference clip, adapt the model to one speaker, or continue from recent speech context.
- Reference conditioning is convenient; adaptation usually improves fidelity but needs more clean data and a separate model or adapter.
- Evaluate intelligibility and speaker similarity separately. A sample can sound natural while sounding like the wrong person.
Connections
- measured_by Speaker Similarity
- part_of Text-to-Speech