Exclusive Visit to a Halted Robot Data Collection Site: Is Embodied AI Recalculating the Data Equation?

On the afternoon of September 1, Radio 42 visited a humanoid robot data training center located in Shijingshan District, Beijing. Although it was a weekday afternoon, the facility was deserted, with no robots operating in any of its zones, including the visitor area, agricultural planting zone, and specialized task zone.

According to reports from multiple media outlets, this humanoid robot data training center has ceased operations, a development that has drawn widespread attention across the industry. Launched in March 2025, the data collection center once deployed over 100 robots and established various realistic scenarios, becoming a benchmark for embodied AI data infrastructure. A relevant party responded that the adjustment was made because centralized remote control data collected in laboratory settings could not meet the needs of real-world applications; the company is now shifting its focus to collecting remote operation data from actual production and living environments.

This reflects a more significant shift within the industry:
Embodied AI still suffers from a lack of data, but what the industry requires is not merely more robots, more workstations, or more teleoperators—but rather a data infrastructure capable of meeting the demands for high quality, large scale, and diversity.
To understand how this infrastructure should be built, we must first address a more fundamental question: Why can't the real-machine data once held in such high hopes support scaling?
Real-world data matters, but it can't scale
The value of teleoperated real-robot data is unquestionable.
When robots complete tasks in real-world environments, they can simultaneously record visual information, action commands, robot states, end-effector feedback, and actual contact results, aligning directly with the action space of specific embodiments.
Therefore, in stages such as post-training for skills, embodiment adaptation, deployment verification, and edge-case collection, real-robot data remains difficult to fully replace with other types of data.
The problem lies in the fact that as embodied AI models pursue larger data scales and broader task coverage, the cost and efficiency bottlenecks of relying solely on real-robot collection become increasingly apparent.

In 2024, DROID, jointly built by Stanford, Berkeley, and other institutions, was one of the representative open real-robot operation datasets globally that year. The project organized continuous data collection across multiple regions in Europe, America, and Asia over 12 months, ultimately producing only about 350 hours of interaction data. The project team noted at the time that multi-environment real-robot collection faces multiple constraints, including logistics, safety, hardware, and human resource investment.
As early as December 2023, Google DeepMind's Open X-Embodiment further aggregated data from 34 laboratories, 60 datasets, and 22 types of robot embodiments worldwide, resulting in just 1 million real-robot trajectories—despite the significant investment of manpower and time, the obtained data remained limited.
Secondly, a more critical bottleneck lies in the fact that these centralized data training centers operate in limited environments and focus on only a few easily collectible tasks, resulting in data with limited value and insufficient generalization, failing to meet the effective application requirements of robots in real-world scenarios.
These projects have demonstrated the value of cross-environment, cross-embodiment real-machine data, while also indicating that the scale expansion of real robot data is far more complex than internet text, images, or even human videos.
Another issue stems from data distribution.
While centralized training centers can improve the utilization rates of both robots and data collectors, the range of scenes and tasks they can construct remains limited. If large volumes of data are concentrated on easily replicable tasks such as desktop organizing, grasping, and moving, the marginal value of newly acquired data may gradually diminish as the collection scale expands.
When robots truly enter factories, warehouses, homes, and even outdoor environments, they often face more complex environmental variations, long-tail situations, and real operational workflows.
This is also a key reason why some recent data collection projects have begun shifting from centralized laboratory collection to real production environments.
Therefore, a more accurate assessment is not that "real-machine data cannot support Embodied AI," but rather that real-machine data alone struggles to bear the entire data scale required for foundation models.
In the future, the physical world experiences learned by robot models will likely come from multiple sources simultaneously, including human videos, simulation environments, real-machine teleoperation, and feedback from actual deployments. There is no simple substitutability among different types of data.
Real-world data more closely reflects the actual working conditions of target robots, offering high value density. However, it is produced slowly and at higher cost, and is constrained by specific hardware forms. In contrast, human-collected and simulation data are easier to scale in volume and task coverage.
Embodied AI data is gradually forming a layered structure: the base consists of first-person human data, the middle layer comprises simulation data, and the apex is real-robot data.
Ego Data Becomes a New Direction for Scaling
In this data structure, first-person human data has become one of the most prominent directions over the past two years, with reasons that are not complex.
Unlike robots, humans can naturally enter homes, malls, warehouses, factories, and outdoor environments to perform fine manipulation, long-horizon tasks, and handle anomalies without needing to deploy specific robots for each demonstration.
Therefore, for the same collection cost, human behavior data can theoretically cover more environments and tasks.
According to media reports from Wall Street View (Huaqiao Jianwen), ego-data (first-person videos and general operation interface data) currently accounts for 40%–50% of total industry data collection time, while real-robot teleoperation data accounts for approximately 30%, with the remainder being simulation and deployment data.
Dr. Xie Chen, Founder and CEO of Guanglun Intelligence, offered a sharper judgment on the future end-state of data in a public speech: Real-robot data will account for only 0.1%, while the remaining 99.9% will consist of these two categories: simulation data and human data.
As the foundational layer, human first-person perspective data can achieve scale precisely because its collection does not rely on fixed embodiments.
The accumulation of this "ego" (first-person perspective) wave has actually been underway for some time.

As early as 2022, Ego4D, built by institutions including Meta, contained 3,670 hours of first-person videos collected by 923 participants across 9 countries and 74 locations, with a scale and scenario richness far exceeding most open real-robot datasets. In February 2026, NVIDIA's EgoScale research further trained robot models using 20,854 hours of action-annotated first-person human videos. The study observed that as the volume of human data expanded, model validation loss improved in an approximately log-linear manner.

This signal was subsequently validated more strongly. In April 2026, Xu Danfei, a student of Fei-Fei Li, led a joint release from Georgia Tech, Stanford, and other institutions: EgoVerse. It contains 1,362 hours of human demonstration data, approximately 80,000 episodes, 1,965 tasks, 240 scenarios, and 2,087 collectors. Validated through multi-robot joint training, it achieved relative improvements of up to 30% across different robot embodiments.

More notably, within EgoVerse, Light Wheel Intelligence represents the standard for scalable, high-quality human data, serving as the sole Chinese enterprise member of its Human Data Committee.
These studies also demonstrate one thing: human data does not naturally translate into robot capabilities.
The physical structures, action spaces, viewpoints, and dynamic conditions of humans and robots still exhibit significant differences. For certain tasks and embodiments, introducing human data may even fail to yield improvements.
Therefore, a more reasonable technical path at present is to use large-scale human data to learn behavioral priors and task structures in the real world, and then employ a substantial amount of robot data to achieve embodiment alignment—aligning the embodiment from “how humans do it” to “how robots do it.”
In other words, the rise of Ego data does not negate the value of real-robot data; rather, it changes the division of labor among different types of data within the overall training framework.
If we simply conceptualize embodied data as a pyramid, the base could consist of human data covering vast real-world environments and human behaviors; the middle layer would comprise simulation data capable of generating tasks at scale, controlling physical parameters, and conducting repetitive trials; while the portion closer to deployment requires real-robot data for embodiment adaptation, post-task training, and deployment verification.
The value of these three types of data cannot be simply compared by “hours,” but they are increasingly taking on distinct roles.
Ego Will Not Replace Real Robots: Data Production Is Increasing Reusability
Understanding this division of labor clarifies the significance of “embodiment-agnostic” approaches. It primarily represents a shift in collection methods.
In traditional real-machine data collection, replacing a robotic arm, dexterous hand, sensor, or control interface often means re-tuning equipment, adapting teleoperation systems, and even re-collecting data. First-person human data decouples behavior collection from specific robots, allowing the same human experience to serve different models, embodiments, and tasks.
But first-person video is not inherently trainable; it requires a 'standardized processing' pathway.
Taking Guanglun Intelligence in EgoVerse as an example, they have built a quality control system covering the entire process of 'edge-side collection—cloud processing—simulation verification'. Thanks to this system, their data pipeline can adapt to diverse hardware. Raw data collected by various multimodal devices, whether Aria, iPhone, head-mounted cameras, or others, can all be fed into the same cloud pipeline. After undergoing VIO (Visual-Inertial Odometry), 3D hand and body annotation, and high-quality action semantic annotation, it is uniformly converted into training assets that meet the same quality standards.
After the data processing stage, Guanglun Intelligence further utilizes its fully self-developed simulation platform SimFoundry and large-scale evaluation platform RoboFinals to verify whether motion trajectories, physical relationships, and task processes can be reasonably reproduced, thereby assessing whether the data truly holds value for robot learning. This is supported by a continuously operating closed-loop evaluation mechanism.
Consequently, the collection end is no longer locked to a single robot, and the data foundation and pre-training capabilities can be reused across platforms.
This is precisely the true watershed for embodied AI data.
The traditional model expands according to workstations, the number of robots, and person-days of collection; the larger the order, the more equipment and personnel costs typically increase in tandem. In contrast, embodiment-free human data can expand scenario and task coverage through distributed collection, then complete annotation, quality inspection, and delivery via unified processing.
As of May 2026, international think tank Interact Analysis tracked at least 90 humanoid robot data collection and training centers in China that are already operational, planned, or under construction, with 64 already operational and 13 deploying over 100 robots. The agency also observed that before 2026, remote-controlled real machines were the mainstream mode; after 2026, new pathways such as first-person human data began to enter data collection centers.
This means data collection sites will not disappear, but their roles will change: from closed, centralized production centers to nodes for real-world scenarios, embodiment alignment, and the recovery of high-value failure data.
What is truly being replaced is the one-off business model of project-based delivery: acquiring a new customer requires purchasing a new robot body, building a new scenario, hiring data collectors, defining tasks, and organizing delivery; after the project ends, the data, equipment, and processes are difficult to reuse. Although it may appear to have many robots and workstations, it is essentially still a labor-intensive, project-based delivery model.
Therefore, what is being replaced is not physical robot data itself, but this 'non-standard, one-off data collection business'—a production method where data can only be tied to a single robot, a specific site, and a one-time output for a single customer—is being replaced by reusable 'standardized data products'.
Building Standardized Data Products and Defining Industry Standards
Since standardized data products are meant to replace non-standard businesses, the question becomes:
What constitutes a standardized product, and how should it be built?
True standardized data products require stable acquisition protocols, task definitions, metadata structures, annotation depth, quality control, delivery formats, and licensing conditions. The same type of product should be identifiable, purchasable, and acceptable to different customers without having to redefine a set of rules for each project.

Guanglun Intelligence’s recent release of EgoSuite-Open100K, the world’s first open-source dataset featuring 100,000 hours of full-modal human behavior data, serves as a typical benchmark for standardized body-centric data products. According to publicly available information on Hugging Face, the project is planned to encompass 100,000 hours of first-person human data, covering over 15,000 tasks, more than 15,000 real-world collection scenarios, seven environmental types, and 128 scenario categories. The initial batch of 10,000 hours has already been launched, with the remainder to be released in phases.
What deserves even greater attention is not merely the singular figure of "100,000 hours," but its data product structure, which meets industry demands for high quality, large scale, and diversity.
EgoStandard is based on leading first-person views and hand pose data, with some subsets adding full-body pose information;
EgoPro adds wrist-view perspectives to supplement information on occlusions, contacts, and fine manipulation; some data provides event-level semantics and supports both LeRobot v3 and MCAP formats.
EgoDemo, designed for developers, offers 50 hours of high-quality samples covering four annotation configurations and two raw video configurations. The data card publicly discloses directory structures, coordinate formats, versions, source identifiers, and verification information.
The dataset employs a license that permits academic research and commercial training while restricting data resale. It resembles a form of controlled openness: lowering industry access costs through public samples, stable interfaces, and community feedback, while preserving reasonable rights.
Guanglun Intelligence summarizes its methodology as 'creating standards, setting benchmarks, and evolving benchmarks': first unifying production, annotation, and delivery rules; then establishing procurement and acceptance scales through reproducible, comparable evaluations; and finally allowing model evaluation and deployment feedback to determine the next round of data production—"without standards, standardized products cannot be resold; without standardized resale, scalability cannot be achieved."
The demand side is validating this logic. According to information disclosed by Guanglun Intelligence, client data demand this year has reached 100 to 1,000 times last year’s levels, jumping from hundreds or thousands of hours to hundreds of thousands or even millions of hours. The company has served over 30 leading model enterprises, embodied AI robot companies, and major internet firms, with each client’s data demand exceeding 200,000 hours. Additionally, it has proposed a five-year, 10-billion-hour joint construction plan for embodied AI data, covering human behavior data, simulated synthetic data, and real-machine deployment feedback data.
Therefore, Guanglun Intelligence’s open-source dataset is not only about building an open ecosystem, but also a verification of standards: only when different teams truly download, train, compare, and provide feedback will the factual standards embedded in standardized data become a common language for cross-client reuse.
To date, Guanglun Intelligence has led or participated in the drafting of more than 20 national and industry-related standards, covering key areas such as data quality, simulation platforms, and high-quality industrial datasets. It has also partnered with the National Robotics Inspection and Certification Center (Headquarters) to jointly promote the construction of key infrastructure for Embodied AI and industrial-grade testing and evaluation standards.
In Conclusion
The shutdown of a single data collection center does not prove that real-world data is invalid. On the contrary, as robots enter factories, warehouses, and homes, post-training data tailored to specific embodiments, real deployment data, and high-value failure cases will become even more important.
However, future data systems will form clearer divisions of labor: high-quality, large-scale, diverse human data will form the foundation of the training tower; simulations will handle scaled trial-and-error and evaluation; and real-robot data will accomplish embodiment alignment, deployment validation, and failure feedback.
Companies like Guanglun Intelligence, which hold both standards and standardized products, are connecting these three elements, enabling experience from one project to be distilled into reusable data products for the next, becoming the most fundamental infrastructure of the industry and supporting the entire open data ecosystem.
The phase of competing on the number of robots at data collection centers is coming to an end.
In the next stage, what will determine an industry position is who can continuously produce reusable data, who can prove data’s effectiveness for models, and who can connect data, evaluation, and deployment feedback into a continuously operating infrastructure.
