Ant Lingbo Tech Launches Industry's First Embodied Native World Action Model LingBot-VA 2.0
Published · Jul 10 · Fri Source · 蚂蚁灵波科技 (CN)

Ant Lingbo Tech Launches Industry's First Embodied Native World Action Model LingBot-VA 2.0

Ant Lingbo Tech launches the industry's first embodied native world action model, LingBot-VA 2.0. The model is pre-trained from scratch based on an autoregressive architecture and adopts four core designs: semantic visual-action tokenizer, causal pre-training paradigm, MoE architecture, and asynchronous inference mechanism. It achieves 150Hz real-time inference on a single card, enabling robots with general control capabilities to deduce and act simultaneously.

KeywordsAntLingboTechLaunchesIndustryFirstEmbodiedNative

The industry's first embodied native world action model is here! Ant Lingbo releases LingBot-VA 2.0

On July 10, Ant Lingbo released the industry's first embodied native world action model, LingBot-VA 2.0. The release marks a key shift in robot foundation models from "built based on digital world models" to "natively designed for the physical world." It represents a key route choice for embodied intelligence development: the robot "brain" no longer relies on the "grafting" of digital world model capabilities, but is natively designed starting from original needs for interacting with the environment, such as dynamic modeling, causal prediction, and real-time execution.

Thanks to the embodied native architecture, LingBot-VA 2.0 demonstrated excellent execution speed and generalization capabilities in real-machine testing. Taking the video below as an example, without relying on any external shooting equipment, the robot can complete multiple rounds of random sparring with humans.

Since the beginning of this year, how world models and embodied intelligence integrate has been a focus of attention for all parties. Starting from the end, based on the "control execution" needs of the physical world, requires continuous "prediction capabilities" that conform to causal laws. Robots face a continuously changing real world; they must not only react to the current situation but also understand what environmental changes an action will trigger and decide the next action accordingly. Current mainstream industry routes mostly rely on video generation models oriented towards digital content creation, then adapt to robot control tasks through fine-tuning.

However, content creation and robot control have different starting points. Content creation cares more about image quality and creativity, while robot control cares more about execution efficiency and the reasonableness of predictions. These differences lead to different capability focuses between digital world video models and physical world video action models from the initial design. Forcibly "fine-tuning" the former to adapt to the latter brings side effects such as knowledge forgetting and reduced generalization.

LingBot-VA 2.0 chooses to face the problem directly and explore a more difficult path—pre-training from scratch based on an autoregressive architecture, building a native foundation model through four core designs.

First, the model introduces a semantic visual-action tokenizer as a new visual encoder, adding alignment of semantic and action information during the visual compression process, making it easier for the model to convert "understanding instructions" into "completing actions" in subsequent training, helping with instruction following and improving action accuracy. Second, the model adopts a strict causal pre-training paradigm, allowing the model to use an autoregressive architecture from the beginning of training, ensuring visual prediction and action generation fully follow a unidirectional time order. Third, the MoE architecture is introduced, effectively expanding model capacity without sacrificing inference efficiency, achieving a balance between performance and efficiency. Finally, real-time closed-loop control is achieved through an enhanced asynchronous inference mechanism, predicting future states while the robot executes actions, and continuously correcting the next decision using the latest real observations. Based on these designs, regarding the industry-wide issue of low execution efficiency in embodied world models, LingBot-VA 2.0 provides an answer of 150Hz real-time inference efficiency on a single card.

From the perspective of "working," robots need to "see clearer," "think clearer," and "act smoother." This week, Ant Lingbo has continuously released and open-sourced multiple models, including: LingBot-Vision and LingBot-Depth 2.0 for spatial perception, LingBot-VLA 2.0 for "one brain, multiple machines" action models, LingBot-World 2.0 for real-time interaction, and LingBot-Video for higher inference efficiency video generation foundation models. The above models represent Ant Lingbo's continuous exploration of specific direction capabilities needed for embodied native, while LingBot-VA 2.0, as the culmination, plays the role of the finale, also officially opening a new stage of embodied native.

Ant Lingbo CEO Zhu Xing stated that on one hand, Lingbo will continue to explore new limits of embodied intelligence, and on the other hand, will accelerate the construction of an open technology ecosystem and scenario ecosystem, helping robots accelerate towards industrial scenarios.

It is reported that Ant Lingbo will fully demonstrate the capabilities of the full-stack brain 2.0 landing scenarios during the 2026 World Artificial Intelligence Conference (WAIC) from July 17 to 20. Audiences can go to the Shanghai World Expo Exhibition & Convention Center (Booths H3-B302, H1-C701) for on-site experience.

This article is provided by Ant Lingbo, Quantum Bit authorized reprint, views belong to the original author.

- 2026 World Artificial Intelligence Conference, held in Shanghai July 17-July 20 2026-07-09

- AI is smart enough, what about action? WAIC first night, let's talk about some real judgments for the next step | Event Registration 2026-07-10

- 11th China Aviation Innovation and Entrepreneurship Competition Registration Opens | Entropy Leaps to the Sky, Boundless New Era 2026-07-09

- Praised by UN Agency! Tianli Qiming "AI+Education" Solution Selected for AI for Good 2026-07-09.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.