• Built by stacking STFT windows: time on one axis, frequency on the other, intensity as energy.
  • Treating it as a grid is why CNN-style layers transfer so well from vision to speech.

Connections