
Ant Lingbo Launches Spatial Perception Model LingBot-Depth 2.0
Ant Lingbo Technology launches the spatial perception model LingBot-Depth 2.0, expanding training data from 3 million to 150 million. It achieved first place in 12 out of 16 depth completion benchmark tests, halving indoor depth error. Simultaneously, the visual base model LingBot-Vision is open-sourced, featuring the industry's first boundary structure pre-training, achieving sub-pixel level boundary localization with only 160 million images. LingBot-Depth 2.0 has passed Orbbec certification, and both parties will cooperate to launch SDKs and integrated camera products.
Robot vision welcomes a new breakthrough! Ant Lingbo's spatial perception model LingBot-Depth 2.0 is officially released
On July 7, Lingbo Technology, an embodied intelligence company under Ant Group, released the spatial perception model LingBot-Depth 2.0. Trained on 150 million scale data, it achieves comprehensive upgrades in edge clarity, small object recognition, long-range depth estimation, and robustness in complex scenarios.
LingBot-Depth is Lingbo's self-developed spatial perception model, equivalent to the robot's eyes in the physical world. Version 1.0 solved spatial perception challenges in complex scenarios like transparent and reflective surfaces. Compared to LingBot-Depth 1.0, LingBot-Depth 2.0 expands training data from 3 million to 150 million, with comprehensive performance upgrades: it achieved first place in 12 out of 16 depth completion benchmark tests; in the most difficult indoor large-area depth missing scenarios, depth error is halved compared to the previous generation (RMSE reduced from 0.132 to 0.062); it performs particularly well in scenarios where traditional depth cameras easily fail, such as glass, mirrors, and transparent objects.
This release also synchronously launches the visual base model for LingBot-Depth 2.0 — LingBot-Vision, building a capability chain for robots from "seeing" to "seeing accurately", aiming to address core challenges in robot vision regarding spatial perception, fine recognition, and complex environment adaptation.
(Figure Caption 1: LingBot-Depth 2.0 completes complete, flat three-dimensional structures in difficult scenarios such as mirrors and glass)
The breakthrough progress of LingBot-Depth 2.0 benefits from LingBot-Vision's outstanding visual representation capabilities. As a general visual model, LingBot-Vision is also the industry's first visual base model to take "boundary structure" as a pre-training target, achieving a breakthrough in spatial perception training paradigms. It possesses sub-pixel level boundary localization and spatial structure understanding capabilities, achieving higher precision and more stable spatial perception capabilities.
LingBot-Vision's pre-training corpus is only 160 million images, one order of magnitude smaller than DINOv3, yet depth estimation accuracy is superior to DINOv3; furthermore, LingBot-Vision's determination of object boundaries is stable enough to continuously track object boundaries in videos. LingBot-Vision open-sources 4 versions this time — ViT-G/L/B/S.
It is understood that besides supporting the training of LingBot-Depth 2.0, LingBot-Vision also possesses "one model, multiple uses" general capabilities.
(Figure Caption 2: LingBot-Depth 2.0 performs leadingly in real sensor depth completion tests)
(Figure Caption 3: Compared with mainstream visual base models, LingBot-Vision's recognition of object boundaries and spatial structures is clearer and more stable)
Currently, LingBot-Depth 2.0 has passed the professional certification of Orbbec Depth Vision Lab. Actual scenario tests show that based on chip-level 3D raw data provided by Orbbec's Gemini 330 series binocular 3D cameras, LingBot-Depth 2.0 has significant improvements in edge clarity, object contour completeness, small object recognition, long-range depth estimation, and robustness under complex lighting and material scenarios.
(Figure Caption 4: LingBot Depth 2.0 passed Orbbec Depth Vision Lab professional evaluation, demonstrating extremely high precision and stability in spatial and temporal depth estimation tasks on multiple sensor models)
In terms of commercialization, Ant Lingbo has launched deep cooperation with Orbbec in many aspects. It is understood that in Orbbec's latest launched body-less data collection product matrix, the RGB-D version of the EGO device will adapt the LingBot-Depth version optimized by Lingbo Technology specifically for data collection scenarios. Subsequently, higher-level commercial version models will be further integrated, continuously completing depth missing, optimizing object edges and spatial structure details, providing a more precise, stable, and usable real-world data base for embodied intelligence model training.
In addition, Orbbec will launch SDK products integrating the latest model capabilities of LingBot-Depth for robot customers to use on the edge, allowing robots using Gemini 330 series cameras to obtain better depth effects; and plans to launch integrated camera products integrating the commercial version of LingBot-Depth by the end of the year, achieving integrated delivery of "3D camera + spatial perception capability". With the release of the two models, cooperation between the two parties is expected to extend to more fields.
Currently, the technical reports of the two models and the model weights of LingBot-Vision have been open-sourced. Ant Lingbo Technology stated that it hopes to build a robot vision base with the industry in an open way, allowing robots to break through the industry bottleneck of "seeing clearly, seeing accurately, seeing steadily" in the real physical world, and accelerating the scaled landing of the embodied industry.
This article is provided by Ant Lingbo, authorized for reprint by Quantum Bit, views belong to the original author.
- 2026 World Artificial Intelligence Conference, held in Shanghai July 17-July 20 2026-07-09
- AI is smart enough, what about action? WAIC first night, let's talk about some real judgments on the next step | Event Registration 2026-07-10
- Industry's first embodied native world action model is here! Ant Lingbo releases LingBot-VA 2.0 2026-07-10
- 11th China Aviation Innovation and Entrepreneurship Competition registration opens | Entropy Leaps to the Sky, Boundless New Era 2026-07-09.
This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.