• Extracts high-level acoustic/semantic features. Often a Conformer or Whisper-style encoder, sometimes with aggressive temporal downsampling.

Connections

No outgoing connections.