• Built around a FastConformer encoder for transcription and speech translation.
  • Aggressive temporal downsampling is the key idea: fewer audio frames while keeping enough local and global context.
  • Task tokens select transcription, translation, punctuation and capitalization behaviour.

Connections