Zhipu Launches GLM-5.1 High-Speed API, GLM-5.1-highspeed
Zhipu Open Platform releases the GLM-5.1 High-Speed API, GLM-5.1-highspeed, with model output speed reaching 400 tokens/s, setting a new global upper limit for large model API speed. GLM-5.1-highspeed is jointly launched by Zhipu and the TileRT team, achieving both flagship-level capabilities and ultra-low latency for the first time in domestic large models. It is suitable for scenarios with extremely high latency requirements such as AI programming, real-time interaction, and real-time voice, and is now open to selected enterprise customers.
Zhipu Open Platform releases the GLM-5.1 High-Speed API, GLM-5.1-highspeed, with model output speed reaching 400 tokens/s, setting a new global upper limit for large model API speed. GLM-5.1-highspeed is jointly launched by Zhipu and the TileRT team, achieving both flagship-level capabilities and ultra-low latency for the first time in domestic large models. It is suitable for scenarios with extremely high latency requirements such as AI programming, real-time interaction, and real-time voice, and is now open to selected enterprise customers.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.