Ant Lingbo Open Sources LingBot-Video, a Video Generation Foundation Model for Embodied Intelligence
Published · Jul 9 · Thu Source · 蚂蚁灵波科技 (CN)

Ant Lingbo Open Sources LingBot-Video, a Video Generation Foundation Model for Embodied Intelligence

Ant Lingbo Technology has open-sourced LingBot-Video, the world's first video generation foundation model based on MoE architecture for embodied intelligence. The model has 30B total parameters, activating only about 3B during inference, achieving 3x efficiency compared to equivalent Dense architectures. The team built a 70,000-hour embodied dataset and introduced a multi-dimensional reinforcement learning reward system to enhance physical plausibility and task completion.

KeywordsAntLingboOpenSourcesLingBot-VideoVideoGenerationFoundation

IT Home July 9 news, today Ant Lingbo Technology officially open-sourced LingBot-Video. It is the world's first open-source video generation foundation model based on Mixture-of-Experts (MoE) architecture for embodied intelligence.

IT Home attaches the official detailed introduction as follows:

Over the past few years, video generation models have made rapid progress in image quality, smoothness, and creative expression. However, for embodied intelligence, a video that looks realistic and has smooth motion may not reflect real physical laws, often making it difficult to support robots in continuous prediction, planning, and task execution. At the same time, embodied intelligence requires models to have higher inference efficiency to adapt to real-time interaction and control loops.

This has led to two different evolutionary directions for video generation: one leading to cinema, serving content creation; the other leading to robots, serving understanding, prediction, and interaction in the physical world. LingBot-Video is our important exploration in opening a new route for video generation for embodied intelligence — redesigning the video pre-training paradigm around the core needs of embodied intelligence, achieving systematic improvements in inference efficiency, physical plausibility, motion understanding, and task completion.

Systematic innovations in architecture, data, and training

For embodied intelligence, LingBot-Video has made systematic innovations in three aspects: architecture, data, and training:

- Architecture: DiT + MoE Design

We replaced the traditional Dense architecture with MoE, controlling single inference costs while expanding model capacity. LingBot-Video's 30B total parameter model activates only about 3B parameters during generation, offering about 3x inference efficiency compared to Dense architectures of the same parameter scale. This design allows the model to gain visual expression capabilities from large-scale parameters while being more suitable for the high-efficiency inference requirements of embodied intelligence.

- Data: Data Profiling Engine and 70,000 Hours of Embodied Data

We built a data profiling engine, introducing robot-related data such as VLA, VLN, and Ego on top of massive internet videos. This data covers scenarios such as dexterous manipulation, robot mobility, and first-person interaction, totaling 70,000 hours. This data helps the model learn the relationship between actions and environmental changes, rather than just learning the surface texture and visual style of videos.

- Training: Multi-dimensional Reinforcement Learning Reward System

We introduced a multi-dimensional reinforcement learning reward system. In addition to conventional metrics such as aesthetics, prompt following, and motion consistency, the model further aligns around physical plausibility and task completion. We also introduced real-world videos as preference signals, making the generated results more consistent with real-world laws and closer to the needs of robots completing tasks in the real world.

Model Performance

On the RBench benchmark jointly released by Peking University and ByteDance, LingBot-Video's total score is 0.620, surpassing Wan2.6 (0.607), Seedance 1.5 Pro (0.584), and Cosmos3 Super (0.581). As a comprehensive evaluation benchmark for robot operation videos, RBench focuses on whether the model can generate robot behaviors consistent with real physical laws. This result indicates that when generating robot-related videos, LingBot-Video better maintains the rationality of the action process and the integrity of task execution.

To further verify LingBot-Video's modeling capability for physical world changes, Ant Lingbo evaluated it from two dimensions: general quality and embodied domain, in an internal benchmark. The results show that LingBot-Video performs better in the embodied domain than major baseline models such as NVIDIA Cosmos 3, Wan 2.2 A14B, LongCat-Video, Hunyuan Video 1.5, and LTX-2.3. In addition, in the Physics-IQ Verified evaluation for physical phenomenon generation and prediction, LingBot-Video also ranked first.

LingBot-Video can be used for robot motion prediction, simulation data generation, motion condition modeling, world model research, and other directions. Currently, LingBot-Video has been officially open-sourced.

Ad Disclaimer: External jump links contained in the text (including but not limited to hyperlinks, QR codes, passwords, etc.) are used to convey more information and save selection time. Results are for reference only. All IT Home articles contain this disclaimer.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.