Extracts high-level acoustic/semantic features. Often a Conformer or Whisper-style encoder, sometimes with aggressive temporal downsampling. Connections No outgoing connections.