Audio quality, conversation quality and business outcome are different evaluations. TTS: intelligibility, naturalness, speaker similarity and first-audio latency. STT: WER by accent, language, noise and domain vocabulary — not one overall number. Agent: task completion, tool correctness, interruption handling, escalation and end-to-end latency. A polished demo can hide failures because the room, script and caller are controlled. Keep a small set of ugly real-world calls and rerun them on every meaningful change. Connections part_of Voice Agent uses Mean Opinion Score uses PESQ uses Speaker Similarity uses Word Error Rate