StepFun Launches New Generation Voice Generation Model StepAudio 2.5 TTS
StepFun officially launches the new generation voice generation model StepAudio 2.5 TTS, featuring three core capabilities: global context control, in-text context control, and zero-shot cloning. Users can precisely control voice details such as emotional tone, rhythm, pauses, and stress through natural language, achieving a leap from "reproducing voice" to "creating expression". The model supports zero-shot cloning of any voice timbre, generating high-quality speech without retraining.
StepFun officially launches the new generation voice generation model StepAudio 2.5 TTS, featuring three core capabilities: global context control, in-text context control, and zero-shot cloning. Users can precisely control voice details such as emotional tone, rhythm, pauses, and stress through natural language, achieving a leap from "reproducing voice" to "creating expression". The model supports zero-shot cloning of any voice timbre, generating high-quality speech without retraining.
A real-time voice large model that truly possesses a "human-like feel". Creating exclusive personas across all dimensions, staying in character even with every breath and chuckle. Inheriting the expressiveness of StepAudio 2.5 TTS and combining industry-leading paralinguistic perception, it instantly understands hesitation and chuckles in tone, outputting matching... at high speed...
May 22, 2026 • StepAudio scored 82.18 on paralinguistic comprehension, demonstrating precise perception of vocal speed, emotion, age, and other acoustic features. On the spoken QA benchmark...
May 22, 2026 • StepAudio-2.5-Realtime is StepFun's new generation real-time voice conversation large model. Starting from "whether it feels like a real person", it reconstructs the warmth and density of voice interaction. It understands the hesitation in your tone, can precisely handle stress and chuckles, and will at the appropriate time...
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.