Built by stacking STFT windows: time on one axis, frequency on the other, intensity as energy. Treating it as a grid is why CNN-style layers transfer so well from vision to speech. Connections related_to Speech-to-Text