Lab arguing that clean-benchmark ASR is the wrong problem. Real rooms (far-field, reverb, SNR < 5 dB, overlapping speech) break models that look great on Librispeech.
They don’t ship one STT. Separate models for target-speaker extraction, directional enhancement, diarization, SSL, tracking, overlap separation, then jointly trained enhancement + STT.
Streaming-first and edge-oriented: causal, RTF < 1, meant to run on phones/laptops without a network. Use cases: home robots, wake-word in noise, in-cabin automotive.