Kuaishou Open-Sources Self-Developed Multimodal Large Model Keye-VL-2.0-30B-A3B
Kuaishou officially open-sources the multimodal large model Keye-VL-2.0-30B-A3B, introducing the DSA sparse attention mechanism to multimodal understanding scenarios for the first time. It supports 256K ultra-long context and achieves long-video temporal causal reasoning. The model surpasses Gemini 2.5 Pro and Gemini 3 Flash on some metrics in the TimeLens video understanding benchmark, unlocks Agent collaboration capabilities for the first time, and covers complex task scenarios such as code, tool calling, and search.
Kuaishou officially open-sources the multimodal large model Keye-VL-2.0-30B-A3B, introducing the DSA sparse attention mechanism to multimodal understanding scenarios for the first time. It supports 256K ultra-long context and achieves long-video temporal causal reasoning. The model surpasses Gemini 2.5 Pro and Gemini 3 Flash on some metrics in the TimeLens video understanding benchmark, unlocks Agent collaboration capabilities for the first time, and covers complex task scenarios such as code, tool calling, and search.
Meet Keye-VL-2.0-30B-A3B — the latest 30B-class flagship base model in the Keye series, purpose-built to push the frontier of long-video understanding and to unlock the first generation of Agent …
June 2, 2026 Keye-VL-2.0-30B-A3B is the flagship multimodal large language model launched by Kuaishou's Keye team, adopting an innovative 30B parameter MoE (Mixture of Experts) architecture design, deeply optimized for long-video understanding and agent capabilities. This model …
May 27, 2026 Today, Kuaishou officially released the new version of multimodal large model Keye-VL-2.0-30B-A3B. As the latest generation 30B-class main base of the Keye family, Keye-VL-2.0-30B-A3B pioneered the introduction of the DSA (DeepSeek Sparse Attention) mechanism into multi …
Keye-VL-2.0-30B-A3B is Kuaishou's open-sourced self-developed multimodal large model, serving as a 30B-class main base. The model introduces DSA sparse attention to multimodal scenarios for the first time, supports 256K ultra-long context, and achieves millisecond-level temporal reasoning for hour-level videos.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.