- Built around a FastConformer encoder for transcription and speech translation.
- Aggressive temporal downsampling is the key idea: fewer audio frames while keeping enough local and global context.
- Task tokens select transcription, translation, punctuation and capitalization behaviour.
Connections
- part_of Speech-to-Text
- related_to Speech Translation