
As embodied AI booms in China, the Beijing-based company is betting that data infrastructure — not robots — will determine the industry’s next winners.
By Jiang Siyuan
Embodied intelligence is, without question, the hottest sector right now. At July’s World Artificial Intelligence Conference (WAIC) in Shanghai, the embodied AI pavilion was the most crowded area. More profound than dancing robots or coffee-serving demos was a new wave entering 1:1 replica factory lines for material handling, sorting, and assembly.
Beyond the exhibition halls, capital, talent, and big tech are converging. In the first half of 2026, domestic embodied AI and related sectors saw over 300 financing rounds, mobilizing more than 90 billion yuan ($13.3 billion). Job postings grew 15-fold in the first four months of the year. Internet giants, automakers, handset makers, and traditional robotics firms are entering from different angles.
Yet beneath the noise, nearly every founder and investor shares the same concern: data.
Over three years, the bottleneck has shifted — from compute to algorithms to data. For embodied AI, without sufficient data, robots cannot learn new skills, adapt to new environments, let alone achieve genuine generalization. The crux: training robots demands data that cannot be generated like internet content, and any single method has inherent limits.
So, starting last year, new data-centric infrastructure has emerged. Some use world models to reduce real-data needs; others mass-produce training data through simulation; some build integrated collection, management, and validation systems.
A fixation on infrastructure
WuWen AI has forged a different path. It has built a complete infrastructure — from real-world collection, cleaning, automated labeling and quality assessment, to synthetic generation, simulation, evaluation and closed-loop validation — linking fragmented stages into a continuously iterating system.
“WuWen AI is the industry’s first to propose and implement a world-model-driven data infrastructure for embodied intelligence,” says founder and CEO Liu Shengxiang.
This fixation on infrastructure is no accident. As a core early leader at Baidu‘s autonomous-driving unit, Liu built out its data and testing framework. When he founded WuWen AI in 2022, he extended that logic to embodied AI — a world-model-driven physical AI data foundation.
Unlike large language models, robots cannot inherit the internet’s text, images and video. They must learn a different world: friction, force variations, causal relationships between actions and feedback, and physical interactions. None of this exists in any ready-made internet.
Globally, high-quality operational data for embodied training is still only a few hundred thousand hours — far from the tens of millions needed. The gap remains an order of magnitude.
Understanding the physical world
Today, most robots rely on imitation learning from human demonstrations. This works for specific tasks but collapses when environments change — shelf height, lighting, object placement. That is why robots perform stably only in narrow settings.
What we should expect is not reciting answers but genuinely understanding the physical world.
Three main data-collection approaches emerged early. Teleoperation yields high precision but is costly and hard to scale. First-person video lacks control, force, and trajectory data. Pure simulation generates massive data cheaply but suffers the Reality Gap — discrepancies in friction, lighting, sensor noise — so virtual skills may not transfer.
After years of experimentation, no single solution suffices. Most companies now hybridize.
Liu sees a robotics-specific scaling law. Real-world “passive wild capture” accumulates physical interaction data; generative world models expand scenarios in simulation. Real and synthetic data complement each other and flow back into training — a closed-loop system of virtual-real integration.
More importantly, this infrastructure could shift robots from imitation learning to reinforcement learning. A high-precision world simulator would let robots undergo trial-and-error millions of times at zero cost, autonomously finding better strategies. Data infrastructure will evolve from a data-supply platform into a training and development foundation for physical AI.
The company’s strategy has drawn comparisons with World Labs, the startup founded by computer scientist Fei-Fei Li, which recently acquired robotics simulation company SceniX as it builds a workflow linking real-world data, simulation and real-world validation.
Liu sees both companies pursuing the same objective: transforming world models from content-generation tools into infrastructure for physical AI.
World Labs recently completed a $1 billion funding round at a $5 billion valuation — signaling capital’s belief in physical-AI infrastructure. “WuWen AI needs to prove that China can produce a world-class infrastructure company in this new era,” Liu says.
Lessons from autonomous driving
Data is only the first step. Whether data is useful only becomes clear after training and real-robot verification. Data collection, screening, labeling, synthesis, training, simulation, verification, deployment and feedback must all connect. Without a feedback mechanism, companies cannot tell which data generated value.
This engineering capability remains scarce — but there is a precedent: autonomous driving, which went through this a decade ago.
Liu says the situation of embodied AI today is where autonomous driving was in 2015–2016. Early competition centered on algorithms, but as models improved, further gains became data dependent. The ability to build a data closed loop — automatically spotting problems, screening samples, annotating, retraining, simulating — became the core competitive capability. Waymo, Cruise and Baidu Apollo all embraced this.
During his Baidu years, Liu’s team built exactly such a system. So when he entered embodied AI, he approached it not as a new algorithm problem, but as a more complex data-closed-loop problem.
He acknowledges robots face far greater complexity than vehicles — they need to understand 3D space, control dozens of degrees of freedom, and handle physical interactions. But the engineering methodology from autonomous driving is what matters: quantifiable, reproducible, continuously verifiable evaluation.
“WuWen AI insists on placing the evaluation system on an equal footing with data and models,” Liu says.
Building the roads rather than the cars
In early-stage tech waves, industrial division of labor cannot form quickly. Many companies must simultaneously build hardware, models, data and application capabilities — not from expansion ambition, but from necessity.
But as the industry moves from demos to scaled R&D, building everything from scratch becomes inefficient. Liu argues the embodied AI industry will eventually develop a clearer division of labor. Rather than every company building robots, foundation models, simulation systems and data pipelines independently, specialized infrastructure providers could handle shared capabilities while robot developers focus on differentiated products.
He compares the model to road construction.
Once roads exist, he said, both Tesla and BYD can use them without each company building its own transportation network. Similarly, robot makers may compete through hardware, algorithms and applications while relying on common infrastructure for data collection, simulation and evaluation.
“The hallmark of a mature industry is not that every company does everything,” Liu says. “It’s that every company becomes exceptionally good at what it does best.”
He has a clear vision: “Embodied intelligence will converge to five to eight general-purpose brain companies, thousands of niche application companies, and a universal data foundation serving all deployment players.”
In a gold rush, eyes are naturally drawn to those wielding shovels; yet, what truly sustains it is not just the prospectors, but also the shovel sellers and the road builders. It is the latter role that Wuwen Zhike aspires to fill.
Source:
LatePost