- https://huggingface.co/BreezeBlue/Breeze-TTS-2
- ~3B open-weight bilingual (EN/ZH) TTS for real-time interaction. Voice clone from a ref clip, voice design from a text description, voice direction to keep identity while steering emotion/pace.
- Inline vocal events:
(laugh),(sigh)in English;[笑],[叹气]in Chinese. - Claims <40ms TTFA and ~0.32 RTF on H100 (warmed fast path). ~7.7 GiB GPU for eager inference. Weights are research/non-commercial; Apache 2.0 on the inference code.
Connections
- part_of Text-to-Speech
- related_to breezeblue
- related_to Voice Cloning