StepFun Launches New Generation Automatic Speech Recognition Model StepAudio 2.5 ASR
StepFun launches the new generation automatic speech recognition model StepAudio 2.5 ASR. The model pioneers the introduction of large language model inference acceleration technology into the speech recognition field. Based on the ASR+MTP-5 architecture, it achieves a 400% increase in inference speed, a 60% reduction in latency, a peak of 500 tokens/s, and an 80% direct cost reduction. It reaches SOTA levels on multiple mainstream Chinese and English evaluation benchmarks. The model reuses a 32K context window and can fully transcribe 30 minutes of long audio in a single pass.
StepFun launches the new generation automatic speech recognition model StepAudio 2.5 ASR. The model pioneers the introduction of large language model inference acceleration technology into the speech recognition field. Based on the ASR+MTP-5 architecture, it achieves a 400% increase in inference speed, a 60% reduction in latency, a peak of 500 tokens/s, and an 80% direct cost reduction. It reaches SOTA levels on multiple mainstream Chinese and English evaluation benchmarks. The model reuses a 32K context window and can fully transcribe 30 minutes of long audio in a single pass.
A real-time speech large model truly possessing a "human-like feel." It creates a dedicated persona across all dimensions, without breaking character even with every breath and chuckle. Inheriting the expressiveness of StepAudio 2.5 TTS, combined with industry-top paralinguistic perception, it instantly understands hesitation and chuckles in tone, outputting rapidly to match...
May 22, 2026 • StepAudio scored 82.18 on paralinguistic comprehension, demonstrating precise perception of vocal speed, emotion, age, and other acoustic features. On the spoken QA benchmark...
May 22, 2026 • StepAudio-2.5-Realtime is StepFun's new generation real-time speech dialogue large model. Starting from "whether it feels like a real person," it reconstructs the warmth and density of speech interaction. It understands the hesitation in your tone, can precisely grasp stress and chuckles, and will at the appropriate time...
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.