Tavus conversational-understanding model. Treats the whole acoustic scene as evidence instead of cancelling noise and waiting for silence.
Jointly models semantics, prosody, speaker identity, backchannels, interruptions, background speech, and its own speaking state (semi-duplex). 10ms frames, ~7ms inference per 80ms of audio.
Can hold 6–8s mid-thought pauses, tell “mhm” from a real barge-in, and ask to repeat when audio is unintelligible. Preliminary TurnBench dev: 92.4% EOT recall, 97.4% interruption recall. GA on Tavus PALs/API.