- Second half of the classic TTS split.
- Autoregressive waveform models such as WaveNet showed how natural neural speech could sound, but one sample at a time was too slow. Later GAN, flow and codec decoders made real-time use practical.
- HiFi-GAN is the useful baseline: fast GAN-based waveform synthesis used behind many acoustic models.
Connections
- part_of Text-to-Speech