SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
TL;DR AI
2 min readKey summary
Researchers introduced SwanData-Speech and SwanVoice, a zero-shot TTS pipeline for 1–4 speakers.
It combines pause-aware text conditioning, a VAE, a flow-matching DiT, and diffusion-based post-training to improve expressive long-form speech.
On SwanBench-Speech, it beat open-source baselines in richness and hierarchy, especially for monologue and dialogue generation.
The system still has room to improve on content accuracy, but it advances coherent, emotionally consistent multi-speaker speech synthesis.
