Common TTS split: text → acoustic representation → waveform. This is the first half. Connections part_of Text-to-Speech