
Emotional TTS Benchmark: Qwen3-TTS, CosyVoice, IndexTTS-2, Fish Audio, and VoxCPM for Japanese and Chinese
Benchmarking five emotional text-to-speech models for Japanese and Chinese across six target emotions, with SenseVoice emotion recognition, emotion2vec anchors, CER, naturalness, runtime, and listening examples.




