Alibaba Tongyi Qianwen Launches Independent AI Voice Input Method "Qianwen Input Method"
Published · Jun 29 · Mon Source · China (CN)

Alibaba Tongyi Qianwen Launches Independent AI Voice Input Method "Qianwen Input Method"

Alibaba Tongyi Qianwen Launches Independent AI Voice Input Method "Qianwen Input Method"

KeywordsAlibabaTongyiQianwenLaunchesIndependentAIVoiceInput

Keyboard input might really be becoming obsolete.

Over the past few years, desktop voice input has been in an extremely awkward position: the system's built-in dictation function is sluggish and full of typos, relegated to being an "accessibility aid"; while third-party input methods have added cloud-based voice recognition, the output text remains unbearable whenever encountering accents, proper nouns, or logically confused long sentences.

Users were forced to struggle between "speaking to input" and "typing to correct," and finally had to honestly type on the keyboard.

But in the last two months, large model technology has reconstructed voice input methods: Alibaba Qianwen launched voice input functionality, Doubao's highly praised mobile voice input was brought to macOS, vertical dark horse Typeless exploded in popularity among independent developers thanks to its Agent capabilities... Even Sogou Input Method has equipped its voice input function with a brand new large model underlying layer.

Is traditional keyboard typing really going to be eliminated? To explore the truth of voice input methods, Lei Keji selected the 4 most mainstream and representative AI-driven voice input products currently on the market, and prepared a horizontal review test for these voice input methods.

Sogou/Doubao/Qianwen/Typeless Compete to be PC "Vibe Mouthpieces"

Before starting the test, let's first introduce the 4 "contestants".

As a veteran national input method, Sogou's latest version on macOS officially introduced Tencent Yuanbao's large model capabilities. In Lei Keji's view, its biggest advantage lies in "seamless transition": Sogou Input Method's AI voice input function is directly integrated into the Sogou Input Method. If you don't use its voice input, the Sogou Input Method is no different from the version you used before.

Alibaba's Qianwen Input Method is not an independent input method, but an independent component within the Qianwen App. It can be used within the Qianwen App, and also outside the Qianwen App, using Qianwen's capabilities to execute voice input. It is worth mentioning that, backed by the Qianwen App, the Qianwen voice input method also possesses Qianwen's corpus regularization, and even summarization and layout capabilities.

Image Source: Qianwen

In comparison, Doubao Input Method is very simple; it is just a "normal" input method with large model voice input capabilities. If you have used Doubao Input Method on your phone, you will definitely not be unfamiliar with the computer version of Doubao Input Method.

As for Typeless, it is a macOS-exclusive voice input tool that has become popular recently in the independent developer circle. It completely abandons the traditional input method concepts of skins and dictionaries, having only a menu bar icon. Its logic is very simple: hold down the shortcut key to speak, release to wait, and the large model will process your recording in the background, outputting text that has been regularized or translated.

Recognition speed varies, but Doubao is surprisingly the best

For voice input, speed determines "whether you want to use it," and accuracy determines "how enjoyable it is to use." To test the recognition accuracy of the four voice input methods, Lei Keji chose a "control variable method": playing a pre-recorded voice clip at a fixed location, and then looking at the speed and accuracy of the four input methods.

First is Sogou Input Method (the following corpus is from Lei Keji's article on the headphone market):

However, from the perspective of overall industry development, Lei Keji believes the transition of the headphone market from incremental to stock will not stop abruptly in 2025. It is certain that, at least in the first half of 2026, the domestic headphone market will still be in this market transition. In Lei Keji's view, 2026 will be the last window period for new audio brands to enter the mainstream market, and AI is the entry ticket for these new headphone forces.

From the recognition effect, Sogou Input Method actually performed quite well, although the sentence segmentation processing had slight flaws. As for the last character "dui" (correct), it was actually because something hit the microphone while I was recording, but this sound was not noise-reduced; instead, it was input as correct text.

In addition, Sogou Input Method has another issue: its voice input preview window is very small, roughly only able to scroll and display less than 10 characters, leaving considerable room for improvement.

Image Source: Lei Keji

Additionally, Sogou voice input recognition speed is quite unstable: sometimes after I finish a sentence, it comes out immediately after two or three seconds; but if it thinks I am speaking a very long text, it must wait until I finish the entire text before starting to output, and this process takes a relatively long time.

Let's look at Qianwen's performance (the following corpus is from Lei Keji's article on the headphone market):

However, from the perspective of overall industry development, Lei Keji believes the transition of the headphone market from incremental to stock will not stop abruptly in 2025. It is certain that, at least until the first half of 2026, the domestic headphone market will still be in this market transition. In Lei Keji's view, 2026 will be the last window period for new audio brands to enter the mainstream market, and AI is the entry ticket for these new headphone forces.

I think Qianwen's voice recognition effect should be discussed in two aspects. First, its voice recognition accuracy is very good, the sentence segmentation is very natural, and you can see it will regularize some of what I say, such as optimizing simple verbal tics or repeated parts. But in terms of recognition speed, if what you say is relatively long, Qianwen's thinking time will also be longer, roughly waiting 3-4 seconds before producing results.

Let's look at Doubao's voice input method (the following corpus is from Lei Keji's article on the headphone market):

However, from the perspective of overall industry development, Lei Keji believes the transition of the headphone market from incremental to stock will not stop abruptly in 2025. It is certain that, at least in the first half of 2026, the domestic headphone market will still be in this market transition. In Lei Keji's view, 2026 will be the last window period for new audio brands to enter the mainstream market, and AI is the entry ticket for these new headphone forces.

Doubao Input Method's working logic is a bit different from the other input methods mentioned earlier; it adopts a real-time transcription mode. As I speak, it transcribes in the foreground simultaneously. This real-time transcription working mode causes Doubao to make some typos when it just starts recognizing.

But because its input is a continuous reasoning process, as long as I continue speaking afterwards, Doubao Input Method will realize the previous errors and automatically correct them before I release my hand to complete the input. Additionally, in terms of recognition speed, Doubao, which has real-time transcription capabilities, is obviously the fastest. The recognition speed is basically only two characters behind my speaking.

Finally, let's look at the performance of the "foreign monk" Typeless (the following corpus is from Lei Keji's article on the headphone market):

However, from the perspective of overall industry development, Lei Keji believes the transition of the headphone market from incremental to stock will not stop abruptly in 2025. It is certain that, at least in the first half of 2026, the domestic headphone market will still be in this market transition. In Lei Keji's view, 2026 will be the last window period for new audio brands to enter the mainstream market, and AI is the entry ticket for these new headphone forces.

In terms of experience, Typeless's performance is a bit like Qianwen; it adopts the mode where I speak first, then it thinks, and then outputs results. It cannot be like Doubao, where it inputs while I speak. So in terms of recognition speed, it is not advantageous like Qianwen.

Typeless's accuracy is acceptable. Like Qianwen, it has voice regularization capabilities, and can directly apply some of my verbal tics or tone words, or parts I modified midway, to the output text, without needing me to modify them repeatedly.

Long text is difficult, is speaking while transcribing a better experience?

Actually, from the tests above, we can also see that because the input modes are different, input methods like Doubao and Sogou that transcribe while speaking, and input methods like Qianwen and Typeless that recognize, think, and process after we finish speaking a complete paragraph, will inevitably have differences in long text recognition.

But the question is, will this difference really affect our daily use? For example, if I speak a long paragraph, will the voice input method overload? Regarding this, we also prepared a long text test.

Because Sogou Input Method uses a voice real-time transcription caching solution, followed by AI polishing of the text. In the long text test, Sogou Input Method did not stutter or show slower recognition speed or longer time consumption because I spoke for one and a half minutes at once. After I finished speaking, the AI polished it for two or three seconds, and then output a complete paragraph of text. I think this is done very well.

As for Qianwen Input Method, limited by the input mode, as long as I keep speaking, Qianwen Input Method will definitely wait until I finish the entire paragraph before processing. Like the short text test, Qianwen's recognition accuracy has no problems, but its recognition and thinking time is significantly longer than the previous short text test. After I finish speaking, it takes about 5-6 seconds to output this paragraph of text all at once.

Doubao Input Method, which transcribes while writing, has better performance in timeliness for long text input. Even if I speak continuously for one minute, it will not overload, and can also achieve the effect where the text appears immediately after I finish speaking.

But Typeless's performance was somewhat unexpected (the following corpus is from Lei Keji's article on magnetic lens reporting):

Of course, any modular solution ultimately cannot avoid ecosystem issues, and magnetic lenses are no exception. In Lei Keji's view, whether magnetic lenses can become a long-term existing product form depends not only on whether the technology is mature, but on whether brands are willing to build a sustainable evolving accessory system around it. In an ideal state, this system may include lens modules of different focal lengths and uses, and even introduce third-party manufacturers to participate.

But from past experience, mobile phone manufacturers are often cautious regarding imaging interfaces and system control rights. Therefore, Lei Keji believes:

For a considerable period of time, magnetic lenses will still exist in a form dominated by manufacturers with limited ecosystems.

It will take on more of an exploration and verification role, rather than rapidly evolving into a universal standard.

But even so, its industry significance still exists. In an imaging market that has already been pushed to the limit by multi-camera algorithms and AI, magnetic lenses at least provide a new problem-solving idea. When body form and module stacking gradually reach their limits, the breakthrough in imaging capabilities may not be within the body.

Although it uses the same record-then-process method as Qianwen, Typeless did not extend its thinking and recognition time because I spoke continuously for 1 minute and 30 seconds. After I finished speaking, I waited less than 2 seconds, and it output the entire paragraph of text, so its efficiency is slightly higher than Qianwen.

But Typeless made the mistake of taking initiative. I just spoke a paragraph, and it automatically formatted the text into an ordered list, very proactively organizing it. I think this is somewhat overstepping.

Mixed Chinese-English speaking and dialects are the ultimate challenge

Obviously, as an input method in the AI era, knowing only Chinese is far from enough. Mixed Chinese-English input, and even dialect input, are the difficulties that test voice input methods. Here, Lei Keji also used the beginning of the article reporting on Google I/O 2026 from a while ago to conduct a simple test on the four input methods.

First is Sogou (the following corpus is from Lei Keji's article on Google I/O reporting):

After much anticipation, it finally arrived. In the early hours of May 20, 2026, Beijing Time, Google I/O 2026 officially opened. Due to the release of new features at Show Event 17, AI became the core topic of this conference. Unlike other AI enterprises, Google simultaneously possesses multiple internet ecosystem entrances such as YouTube, Google Web Search, and Android. Therefore, how to empower the above ecosystems with AI technology became the key topic of this conference.

Although in terms of function, Sogou did not conduct a.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.