Strong accuracy baseline. Vanilla implementation is not built for streaming. Connections part_of Speech-to-Text uses Mel Spectrogram