Alibaba Open-Sources Real-Time Speech Recognition Large Model Fun-ASR-Realtime
Published · Jul 6 · Mon Source · 千问大模型 (CN)

Alibaba Open-Sources Real-Time Speech Recognition Large Model Fun-ASR-Realtime

Alibaba officially launches the upgraded real-time speech recognition large model Fun-ASR-Realtime. First-word latency is controlled at the hundred-millisecond level, and recognition accuracy approaches that of offline models. It supports 16 dialects and 30 languages. The model possesses context understanding capabilities and can self-correct; in 16 dialect recognition tests, the average character accuracy rate is 88.62%, leading related products from Volcano and Tencent.

KeywordsAlibabaOpen-SourcesReal-TimeSpeechRecognitionLargeModelFun-ASR-Realtime

Alibaba officially launches the upgraded real-time speech recognition large model Fun-ASR-Realtime. First-word latency is controlled at the hundred-millisecond level, and recognition accuracy approaches that of offline models. It supports 16 dialects and 30 languages. The model possesses context understanding capabilities and can self-correct; in 16 dialect recognition tests, the average character accuracy rate is 88.62%, leading related products from Volcano and Tencent.

November 7, 2025 • Fun-ASR-RealTime Python SDK, Large Model Service Platform Bailian: This article introduces the parameters and interface details of the Fun-ASR real-time speech recognition Python SDK. Alibaba Cloud Bailian has launched dedicated business spaces for the North China 2 (Beijing) and Singapore regions...

4 days ago • FAQ Q: What model is Fun-ASR-Realtime? A: Fun-ASR-Realtime is a real-time speech recognition model launched by Alibaba Tongyi, mainly used to convert continuous speech streams into text in real time, suitable for real-time subtitles, meeting transcription, voice assistants, customer service...

3 days ago • Intro On July 6, the Alibaba Qwen team upgraded the Fun-ASR-Realtime real-time speech recognition large model. A single model supports 30 languages + 16 dialects, first-word latency is compressed to the hundred-millisecond level, and streaming recognition accuracy has approached offline model levels. At the same time...

4 days ago • According to the introduction, Fun-ASR-Realtime first-word latency is controlled at the hundred-millisecond level, streaming recognition accuracy approaches offline levels, and it simultaneously supports seamless multi-language switching. In addition, the model has undergone specialized optimization for multi-language scenarios in East and Southeast Asia such as Thai, recognition accu...

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.