• Second half of the classic TTS split.
  • Autoregressive waveform models such as WaveNet showed how natural neural speech could sound, but one sample at a time was too slow. Later GAN, flow and codec decoders made real-time use practical.
  • HiFi-GAN is the useful baseline: fast GAN-based waveform synthesis used behind many acoustic models.

Connections