A framework for evaluating multi-turn LLM conversations with support for text, realtime audio, and speech-to-speech models.

check out here