Ant Group's Robotics Unit Launches LingBot-VA 2.0 Embodied AI Model
Ant Group’s robotics subsidiary, Lingbo, has released LingBot‑VA 2.0, the first embodied‑native foundation model designed for robot manipulation. The model replaces conventional video‑action pipelines with a causal diffusion transformer (DiT) that jointly processes video and action streams. It incorporates a mixture‑of‑experts video backbone (about 13 billion parameters, 1.9 billion active) and a dense action expert, enabling faster inference and closed‑loop control. Training integrates multi‑chunk prediction, semantic alignment, and forward dynamics objectives, allowing the system to learn how actions reshape the world directly from unlabeled web video.
Lingbo’s strategy emphasizes software and data over hardware, repurposing Ant Group’s extensive AI infrastructure and transaction‑level data from its Alipay payment platform. By leveraging real‑time decision‑making and large‑scale distributed inference capabilities, the unit aims to build a “robot brain” that can achieve the same reliability in physical environments as Ant’s financial systems. The approach reflects a broader shift in Chinese robotics where the competitive edge is moving from hardware design to intelligent, data‑driven software layers.