NetEase Youdao Open Sources TTS Speech Synthesis Engine Confucius4-TTS
NetEase Youdao open sources the TTS model Confucius4-TTS. The model achieves three major breakthroughs: 3-second zero-shot voice cloning, accent-free cross-lingual synthesis across 14 languages, and emotional prosody transfer. The underlying model adopts an end-to-end architecture of speech encoder + large language model + flow matching generation. The complete 54G weights support local offline deployment.
NetEase Youdao open sources the TTS model Confucius4-TTS. The model achieves three major breakthroughs: 3-second zero-shot voice cloning, accent-free cross-lingual synthesis across 14 languages, and emotional prosody transfer. The underlying model adopts an end-to-end architecture of speech encoder + large language model + flow matching generation. The complete 54G weights support local offline deployment.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.