• https://www.tavus.io/blog/sparrow-2
  • Tavus conversational-understanding model. Treats the whole acoustic scene as evidence instead of cancelling noise and waiting for silence.
  • Jointly models semantics, prosody, speaker identity, backchannels, interruptions, background speech, and its own speaking state (semi-duplex). 10ms frames, ~7ms inference per 80ms of audio.
  • Can hold 6–8s mid-thought pauses, tell “mhm” from a real barge-in, and ask to repeat when audio is unintelligible. Preliminary TurnBench dev: 92.4% EOT recall, 97.4% interruption recall. GA on Tavus PALs/API.

Connections