Zhipu Opensource New Generation Flagship Model GLM-5.1
Published · Apr 8 · Wed Source · 智谱 (CN)

Zhipu Opensource New Generation Flagship Model GLM-5.1

Zhipu opensource flagship model GLM-5.1, the world's strongest open-source model, capable of independent continuous work for over 8 hours, autonomously completing complex engineering tasks. The model's coding ability ranks third globally and first domestically in benchmark tests such as SWE-Bench Pro, surpassing GPT-5.4 and Claude Opus 4.6. Real-world tests show it can build a complete Linux desktop system in 8 hours, optimize vector database performance by nearly 7 times, and iterate on ML loads over 24 hours to achieve 3.6 times acceleration.

KeywordsGPTClaudeZhipuOpensourceNewGenerationFlagshipModel

Overview

GLM-5.1 is a high-performance model built by Zhipu for complex code and long-horizon task scenarios. Coding ability is greatly enhanced, and long-horizon tasks are significantly improved. It can work continuously and autonomously for up to 8 hours in a single task, completing a full closed loop from planning, execution to iterative optimization, delivering engineering-grade results. In terms of comprehensive ability and Coding ability, GLM-5.1 overall performance aligns with Claude Opus 4.6, and demonstrates stronger continuous working ability in long-horizon autonomous execution, complex engineering optimization, and real development scenarios. It is an ideal base for building Autonomous Agents and Long-horizon Coding Agents.

Positioning

High-intelligence base model

Input Modality

Text

Output Modality

Text

Context Window

200K

Max Output Tokens

128K

Capabilities Support

Thinking Mode

Provides multiple thinking modes covering different task requirements

Streaming Output

Supports real-time streaming response, improving user interaction experience

Function Call

Powerful tool calling ability, supporting integration of various external tools

Context Caching

Intelligent caching mechanism, optimizing long conversation performance

Structured Output

Supports structured format output such as JSON, facilitating system integration

MCP

Can flexibly call external MCP tools and data sources, expanding application scenarios

Recommended Scenarios

Agentic Coding

Further optimized for typical Agentic Coding scenarios such as Claude Code and OpenClaw. Possesses stronger long-horizon planning, step-by-step execution, process adjustment, and result delivery capabilities. Performance is significantly improved in long-horizon development tasks and complex programming problems, suitable for real engineering tasks with multiple stages and strong dependencies.

General Conversation

Performs more stably in open-ended Q&A, complex instruction understanding, and multi-turn communication scenarios. Responses are richer in dimensions and more complete in content, with stronger instruction following ability and long-context understanding ability. Suitable for high-quality daily assistants and complex information interaction scenarios.

Creative Writing

Further enhanced in literary expression, plot extension, character portrayal, and language style control. Applicable to writing tasks requiring high expressiveness and consistency, such as novel fragments, story settings, and copywriting creation.

Artifacts / Frontend Development

Suitable for web page, interactive page, and frontend prototype generation scenarios. Generated results further reduce template feel, visual expression is more diverse, and overall completion degree of frontend tasks is higher. Can more efficiently support rapid landing from requirements to usable products.

Office Productivity

Overall improvement in document production tasks such as PPT, Word, PDF, and Excel. Capable of completing more complex content organization, layout design, and structured output. Default aesthetics and finished product quality are significantly enhanced. Suitable for high-intensity production scenarios such as long documents, reports, textbooks, and papers.

Detailed Introduction

Comprehensive & Coding Ability: Aligning with Global Top Level

GLM-5.1 reaches the global first tier in comprehensive and Coding ability, with overall performance aligning with Claude Opus 4.6, and ranking at the forefront in multiple key evaluations. In the SWE-Bench Pro benchmark test, GLM-5.1 achieved a score of 58.4, surpassing GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro, refreshing the global best performance. Meanwhile, across 12 representative benchmarks covering reasoning, programming, Agent, tool calling, and browsing, GLM-5.1 also demonstrates a comprehensive and balanced capability structure. This indicates that GLM-5.1's improvement is not a single-point breakthrough, but a synchronous enhancement in three dimensions: general intelligence, real programming, and complex task execution, making it more suitable as a base model for general Agent systems and engineering production scenarios.

Long-horizon Task Ability: Moving Towards 8-Hour Level Continuous Work

GLM-5.1's long-horizon task (Long Horizon Task) ability is significantly improved, focusing on enhancing the model's continuous execution, closed-loop optimization, and engineering delivery capabilities under complex goals. Compared to models mainly based on minute-level interactions, GLM-5.1 can work continuously and autonomously for up to 8 hours in a single task, completing the full process from planning, execution, testing to fixing and delivery. Under the same evaluation standards, GLM-5.1 is one of the few models with 8-hour level continuous working ability, and a representative of Chinese models to first reach this level. The standard for measuring model ability is evolving from "smarter in single turn" to "how long it can work stably and deliver what in long-horizon tasks". This kind of ability is not just longer context, but requires the model to maintain goal consistency during long-term execution, reducing strategy drift, error accumulation, and invalid trial and error, truly possessing autonomous execution ability for complex engineering tasks.

Engineering Delivery Ability: Evolving from Code Generation to Fully Autonomous Agents

One of the core breakthroughs of GLM-5.1 is forming an autonomous closed loop of "experiment-analyze-optimize" in long-horizon tasks, rather than staying at the level of one-time code generation. The model can actively run benchmarks, identify bottlenecks, adjust strategies, and continuously improve result quality through multiple iterations. In typical cases, GLM-5.1 can build a complete Linux desktop system from scratch within 8 hours; autonomously conduct 655 rounds of iteration to complete the entire optimization link, improving vector database query throughput to 6.9 times that of the initial official version; on the KernelBench Level 3 optimization benchmark, complete thousands of rounds of tool calling to optimize real machine learning model loads, achieving a 3.6 times geometric mean acceleration ratio, far exceeding the 1.49 times of torch.compile max-autotune mode. These results show that GLM-5.1 already possesses the ability to autonomously explore, continuously improve, and stably deliver in complex engineering environments, capable of handling higher value tasks such as system construction, performance optimization, and long-horizon Coding Agents.

Usage Resources

Experience Center

Quickly test the model's effect in business scenarios

API Documentation

API calling methods

Call Examples

The following are complete call examples to help you get started with the GLM-5.1 model quickly. - cURL - Python - Java - Python(Old) - Basic Call - Streaming Call.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.