
Zhipu Opensource New Generation Flagship Model GLM-5.1
Zhipu opensource flagship model GLM-5.1, the world's strongest open-source model, capable of independent continuous work for over 8 hours, autonomously completing complex engineering tasks. The model's coding ability ranks third globally and first domestically in benchmark tests such as SWE-Bench Pro, surpassing GPT-5.4 and Claude Opus 4.6. Real-world tests show it can build a complete Linux desktop system in 8 hours, optimize vector database performance by nearly 7 times, and iterate on ML loads over 24 hours to achieve 3.6 times acceleration.
Overview
GLM-5.1 is a high-performance model built by Zhipu for complex code and long-horizon task scenarios. Coding ability is greatly enhanced, and long-horizon tasks are significantly improved. It can work continuously and autonomously for up to 8 hours in a single task, completing a full closed loop from planning, execution to iterative optimization, delivering engineering-grade results. In terms of comprehensive ability and Coding ability, GLM-5.1 overall performance aligns with Claude Opus 4.6, and demonstrates stronger continuous working ability in long-horizon autonomous execution, complex engineering optimization, and real development scenarios. It is an ideal base for building Autonomous Agents and Long-horizon Coding Agents.
Positioning
High-intelligence base model
Input Modality
Text
Output Modality
Text
Context Window
200K
Max Output Tokens
128K
Capabilities Support
Thinking Mode
Provides multiple thinking modes covering different task requirements
Streaming Output
Supports real-time streaming response, improving user interaction experience
Function Call
Powerful tool calling ability, supporting integration of various external tools
Context Caching
Intelligent caching mechanism, optimizing long conversation performance
Structured Output
Supports structured format output such as JSON, facilitating system integration
MCP
Can flexibly call external MCP tools and data sources, expanding application scenarios
Recommended Scenarios
Agentic Coding
Further optimized for typical Agentic Coding scenarios such as Claude Code and OpenClaw. Possesses stronger long-horizon planning, step-by-step execution, process adjustment, and result delivery capabilities. Performance is significantly improved in long-horizon development tasks and complex programming problems, suitable for real engineering tasks with multiple stages and strong dependencies.
General Conversation
Performs more stably in open-ended Q&A, complex instruction understanding, and multi-turn communication scenarios. Responses are richer in dimensions and more complete in content, with stronger instruction following ability and long-context understanding ability. Suitable for high-quality daily assistants and complex information interaction scenarios.
Creative Writing
Further enhanced in literary expression, plot extension, character portrayal, and language style control. Applicable to writing tasks requiring high expressiveness and consistency, such as novel fragments, story settings, and copywriting creation.
Artifacts / Frontend Development
Suitable for web page, interactive page, and frontend prototype generation scenarios. Generated results further reduce template feel, visual expression is more diverse, and overall completion degree of frontend tasks is higher. Can more efficiently support rapid landing from requirements to usable products.
Office Productivity
Overall improvement in document production tasks such as PPT, Word, PDF, and Excel. Capable of completing more complex content organization, layout design, and structured output. Default aesthetics and finished product quality are significantly enhanced. Suitable for high-intensity production scenarios such as long documents, reports, textbooks, and papers.
Detailed Introduction
Comprehensive & Coding Ability: Aligning with Global Top Level
GLM-5.1 reaches the global first tier in comprehensive and Coding ability, with overall performance aligning with Claude Opus 4.6, and ranking at the forefront in multiple key evaluations. In the SWE-Bench Pro benchmark test, GLM-5.1 achieved a score of 58.4, surpassing GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro, refreshing the global best performance. Meanwhile, across 12 representative benchmarks covering reasoning, programming, Agent, tool calling, and browsing, GLM-5.1 also demonstrates a comprehensive and balanced capability structure. This indicates that GLM-5.1's improvement is not a single-point breakthrough, but a synchronous enhancement in three dimensions: general intelligence, real programming, and complex task execution, making it more suitable as a base model for general Agent systems and engineering production scenarios.
Long-horizon Task Ability: Moving Towards 8-Hour Level Continuous Work
GLM-5.1's long-horizon task (Long Horizon Task) ability is significantly improved, focusing on enhancing the model's continuous execution, closed-loop optimization, and engineering delivery capabilities under complex goals. Compared to models mainly based on minute-level interactions, GLM-5.1 can work continuously and autonomously for up to 8 hours in a single task, completing the full process from planning, execution, testing to fixing and delivery. Under the same evaluation standards, GLM-5.1 is one of the few models with 8-hour level continuous working ability, and a representative of Chinese models to first reach this level. The standard for measuring model ability is evolving from "smarter in single turn" to "how long it can work stably and deliver what in long-horizon tasks". This kind of ability is not just longer context, but requires the model to maintain goal consistency during long-term execution, reducing strategy drift, error accumulation, and invalid trial and error, truly possessing autonomous execution ability for complex engineering tasks.
Engineering Delivery Ability: Evolving from Code Generation to Fully Autonomous Agents
One of the core breakthroughs of GLM-5.1 is forming an autonomous closed loop of "experiment-analyze-optimize" in long-horizon tasks, rather than staying at the level of one-time code generation. The model can actively run benchmarks, identify bottlenecks, adjust strategies, and continuously improve result quality through multiple iterations. In typical cases, GLM-5.1 can build a complete Linux desktop system from scratch within 8 hours; autonomously conduct 655 rounds of iteration to complete the entire optimization link, improving vector database query throughput to 6.9 times that of the initial official version; on the KernelBench Level 3 optimization benchmark, complete thousands of rounds of tool calling to optimize real machine learning model loads, achieving a 3.6 times geometric mean acceleration ratio, far exceeding the 1.49 times of torch.compile max-autotune mode. These results show that GLM-5.1 already possesses the ability to autonomously explore, continuously improve, and stably deliver in complex engineering environments, capable of handling higher value tasks such as system construction, performance optimization, and long-horizon Coding Agents.
Usage Resources
Experience Center
Quickly test the model's effect in business scenarios
API Documentation
API calling methods
Call Examples
The following are complete call examples to help you get started with the GLM-5.1 model quickly. - cURL - Python - Java - Python(Old) - Basic Call - Streaming Call.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.