
ByteDance Launches New Agent Model Seed2.1
ByteDance launches the Seed2.1 series of models, positioned as new agents for real productivity scenarios. The model shows significant improvements in three dimensions: general Agent capabilities, end-to-end code delivery, and multimodal understanding. It leads in multiple benchmarks including Workspace Bench, GDPval, and MobileWorld, with a crowdtest win rate exceeding Claude Opus 4.6 in Coding scenarios.
Seed2.1 Officially Released, Deepening AI Productivity
Date
2026-06-23
Category
Model Release
The Seed model series has always been committed to uncovering real user needs and stimulating user creativity. After the release of Seed2.0, we continued to track user feedback and observed that user expectations for the model further pointed towards more reliable responses and more stable delivery.
Against this background, we are pleased to introduce the Seed2.1 series, a new agent oriented towards real productivity scenarios. Seed2.1 aims to solve complex needs in daily life, professional work, and frontier exploration. It continuously incorporates feedback from internal and external users and developers, and combines real cases to calibrate model optimization directions. In terms of evaluation, we also focus more on the model's performance in actual workflows rather than relying solely on static benchmark scores.
We will introduce the real capabilities of Seed2.1 from the following three dimensions:
- More reliable general Agent capabilities: Seed2.1's general Agent capabilities have significantly improved, further strengthening task delivery capabilities across tools and environments. When facing office tasks with high economic value and complex consultations in personal life, it can stably complete multi-step tasks such as project planning, file processing, and tool calling, producing actionable results.
- More stable code engineering delivery capabilities: Seed2.1 has improved the end-to-end delivery capability of Coding. It can complete tasks such as requirement understanding, function implementation, bug fixing, runtime environment setup, and result verification in real enterprise-level development tasks, forming stable delivery.
- Stronger multimodal and basic capabilities: Seed2.1 has further improved basic capabilities such as multimodal understanding, knowledge, and reasoning. It processes complex visual information and video content more accurately, providing foundational support for Agentic scenarios, code engineering, and frontier exploration.
The Seed2.1 series of models has been launched on the Doubao product and TRAE. At the same time, the API for this series of models has been synchronized to Volcano Engine. Everyone is welcome to experience and provide feedback.
Project Homepage (including Model Card):
https://seed.bytedance.com/seed2_1
Experience Entry:
1) Select "Office Task" mode on Doubao PC version or Doubao App
2) Select Doubao-Seed-2.1-Pro or Doubao-Seed-2.1-Turbo in TRAE Work or TRAE IDE built-in models
3) Select Doubao-Seed-2.1-Pro or Doubao-Seed-2.1-Turbo in Volcano Ark Experience Center
General Agent Capabilities Significantly Improved, Executing Complex Tasks More Reliably
When models enter productivity scenarios, users need not just a single answer, but for the model to continuously advance tasks around a goal and produce usable results. Around this direction, Seed2.1 has further strengthened general Agent capabilities. Whether facing high economic value work tasks or complex consultations for personal life, the model can deliver reliably.
Facing high economic value work tasks, in the past, users might need to consult external consultants or professional service teams to assist in completion; now, the model can participate in data analysis, solution design, content planning, and result organization, helping users advance work that originally required professional support, achieving cost reduction and efficiency increase.
Seed2.1 performs stably on Workspace Bench and Agent Startup Bench benchmarks. Seed2.1 Pro achieved the highest score on the GDPval benchmark. Among them, Workspace Bench focuses on information retrieval, associated understanding, and result generation for complex files in work; Agent Startup Bench comprehensively evaluates the model's answer quality by researching and interviewing real AI-native startups combined with expert opinions; GDPval measures the completion quality and economic value of the model in real-world work tasks. Evaluation results indicate that in AI workflows close to real work tasks, Seed2.1 can establish connections between complex materials and task goals, and produce deliveries with economic benefits.
In addition, on higher difficulty and more professional tasks, Seed2.1 also performs well. Among them, Seed2.1 Pro is in the first tier of current evaluated models in the Agents' Last Exam (ALE) benchmark evaluation, reflecting strong competitiveness in complex professional tasks. It is worth noting that this evaluation was released recently, and models cannot fully optimize for this test in the short term, which can more truly measure the model's generalization ability when facing new task scenarios. This result indicates that the general Agent capabilities possessed by Seed2.1, such as task planning, tool usage, long-term execution, information integration, and result delivery, can be well transferred to previously unseen high-threshold professional workflows.
In the Agents' Last Exam benchmark evaluation, the left side is the complete pass rate, and the right side is the average comprehensive score.
Facing complex consulting scenarios in personal life, the quality and reliability of responses from the Seed2.1 series of models have further improved.
Such needs are often not simple Q&A. Users will simultaneously provide consulting background, past records, industry reports, and various other information. The content is also distributed across different formats such as documents, PDFs, and images, forming a complex consulting scenario requiring comprehensive reasoning, judgment, and suggestions.
Seed2.1 performs stably on benchmarks such as xDailyBench and Doubao Multi-Turn Bench, and maintains competitiveness on benchmarks such as Toolathlon and SeedClawBench. This indicates that in more than 30 vertical scenarios such as daily life and learning research, the model can better understand real user needs, combine user preferences to give high-quality suggestions, and when necessary, call different tools and use appropriate Skills to produce reliable responses.
SeedClawBench is an internal benchmark developed by Seed, used to evaluate the Agent's ability to provide practical assistance in OpenClaw-style, user-oriented scenarios.
Around scenarios such as teaching, general office work, and professional research, Seed2.1 can stably output lesson plan PPTs, complete complex table analysis, and generate industry reports.
In addition, based on stable visual understanding capabilities, Seed2.1 can better process visual information, understand user goals, and advance subsequent execution and delivery in complex tasks. Seed2.1 generally shows strong competitiveness on Visual Agent related benchmarks such as Claw-Eval (MM). This means the model can not only understand complex visual information such as documents, videos, images, and spatial structures, but also organize and analyze visual information around task goals, and form interactive and deliverable Agent results, such as generating a floor plan based on multi-view images, or completing tasks such as information retrieval, content generation, and code writing based on visual information.
Image2FloorPlan is an internally built evaluation set, examining the task of understanding multiple real photos and drawing a floor plan.
In the exploration of professional productivity scenarios, we found that real workflows do not occur in a single fixed interface, but require switching between chat, search, browser, code repositories, files, and external tools. Therefore, Seed2.1 has further optimized towards the general Computer-Use Agent (CUA) direction, allowing the model to more stably advance in tasks across environments, tools, and interaction methods.
Among them, facing mobile GUI tasks, the model needs to understand screen content, judge the next operation, and complete continuous actions such as clicking, inputting, and switching applications. Seed2.1 achieved the highest score in the MobileWorld benchmark, indicating that it can more stably advance operations in mobile phone tasks. At the same time, the model maintains competitiveness on OSWorld, and through reinforcement learning, guides the Agent to naturally switch optimal choices between GUI and non-GUI action spaces, reducing the average number of steps required to complete tasks by 16%, further improving task execution efficiency.
In addition, Seed2.1 also performs outstandingly on the CreativeWork benchmark. This benchmark covers three representative environments: Notion, Canva, and Figma, meaning the model can understand complex goals, decompose execution steps, and autonomously switch between tool calling and GUI interaction in various tasks such as document management, visual design, and interface editing, stably completing tasks.
CreativeWork is a benchmark developed by Seed, used to evaluate the Agent's ability to collaboratively use GUI and MCP tools in real productivity scenarios.
Coding End-to-End Capabilities Significantly Strengthened, Enterprise Production Scenario Delivery Stable
Focusing on the Coding Agent direction, Seed2.1 combines public benchmarks, crowdtest developer feedback, and internal evaluations to comprehensively assess model performance. Among them, public benchmarks mainly focus on the model's capability boundaries in general code tasks, while crowdtest developer feedback better reflects the model's actual value in real engineering scenarios.
In public benchmarks, Seed2.1 Pro maintains competitiveness on the ProgramBench benchmark, indicating the model has the capability to complete system-level engineering from scratch, independently completing software system architecture design and code implementation.
At the same time, Seed2.1 Pro performs well on the NL2Repo-Bench benchmark. This benchmark mainly evaluates the model's ability to convert natural language requirements into repository-level code changes, which is closer to real software engineering scenarios. Evaluation results indicate that Seed2.1 can understand the architecture, dependencies, and business logic of the entire code repository, and perform multi-file collaborative modifications, finally delivering maintainable and runnable engineering code.
In crowdtest developer evaluations, we invited developers to submit engineering tasks based on real code repositories and compare anonymous model outputs. The results show that in tasks closer to real Coding processes, Seed2.1 received higher evaluations on final completion quality. Among them, Seed2.1 Pro achieved a 59.1% win rate compared to Claude Opus 4.6.
In addition, the Seed2.1 Preview version also participated in human preference evaluations for frontend scenarios recently. In the Code Arena: Frontend leaderboard, the model ranked 8th with 1539 points, and entered the top 10 in 5 out of 7 frontend subcategories.
Multimodal Understanding and Basic Capabilities Continue to Lead, Further Serving Agentic Scenarios
Seed2.1 continues to deepen multimodal capabilities. In various visual and video understanding tasks, it achieved SOTA results on multiple evaluation sets, maintaining industry-leading standards, and further serving Agentic scenarios.
Facing visual understanding scenarios, Seed2.1 Pro achieved the highest scores on multiple benchmarks such as CharXiv-RQ and MeasureBench, reflecting further improvements in the model's complex document understanding, chart reading, numerical identification, and visual detail judgment. Such capabilities can help the model reduce misreading when processing PDFs, reports, charts, and multi-page materials, and enhance perception of unstructured information.
Seed2.1 also achieved the best results on the ERQA benchmark, indicating that the model's spatial understanding capability has further strengthened, can better.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.