Often a linear projection or a small Conformer stack. The goal is structural alignment so the LLM can treat audio embeddings like text embeddings. Connections No outgoing connections.