Xiaomi Launches MiMo-V2.5-Pro UltraSpeed Mode
Published · Jun 9 · Tue Source · IT之家 (CN)

Xiaomi Launches MiMo-V2.5-Pro UltraSpeed Mode

Xiaomi, in collaboration with TileRT, has launched the MiMo-V2.5-Pro UltraSpeed mode, achieving a generation speed exceeding 1000 tokens/s for a trillion-parameter model on general-purpose GPUs for the first time. This mode is priced at three times that of the standard version, offers approximately 10 times faster output speed, and is available only via API. The MiMo-V2.5-Pro UltraSpeed mode is available through a limited-time application system, with priority review given to enterprises and professional developers.

KeywordsAPIXiaomiLaunchesMiMo-V2.5-ProUltraSpeedModeTileRTGPUs

Xiaomi, in collaboration with TileRT, has launched the MiMo-V2.5-Pro UltraSpeed mode, achieving a generation speed exceeding 1000 tokens/s for a trillion-parameter model on general-purpose GPUs for the first time. This mode is priced at three times that of the standard version, offers approximately 10 times faster output speed, and is available only via API. The MiMo-V2.5-Pro UltraSpeed mode is available through a limited-time application system, with priority review given to enterprises and professional developers.

June 8, 2026 • MiMo, in collaboration with TileRT, releases the UltraSpeed mode of Xiaomi MiMo-V2.5-Pro — breaking 1000 tokens/s generation speed on a 1T-parameter model for the first time on …

June 8, 2026 • MiMo × TileRT jointly release the UltraSpeed mode of Xiaomi MiMo-V2.5-Pro. Through extreme codesign between the model and system, they have broken through 1000 tokens/s generation speed for a trillion-parameter model on general-purpose GPUs for the first time.

The UltraSpeed experience mode of MiMo-V2.5-Pro — a trillion-parameter (1T) flagship model reaching inference speeds of up to 1000 tokens/s, built for the most demanding real-time scenarios.

June 11, 2026 • Finally got the UltraSpeed mode of Xiaomi MiMo-V2.5-Pro, achieving an output speed of 1000 tokens/s for a trillion-parameter inference model. This means all models in this chart are more than 6 times slower than this model. Also, recently, V2.5-Pro also had a price drop ….

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.