Behind ACE-Data-0: A Conversation with Daxiao Robotics on the Next Step for Embodied Data

Can robots learn everything just by watching videos? If so, you're overestimating them.

Consider the simplest everyday scenario: picking up a water cup from a table. For humans, this requires almost no thought; for robots, it is a long chain involving perception, movement, grasping, contact, changes in object state, and task memory.

Videos can tell a model to 'pick up the cup,' but they don't reveal when fingers make contact with the cup's surface, how grip force changes, how the cup moves in 3D space, or what occluded fingers are doing. The model knows none of this.

This highlights the data dilemma in Embodied AI: while there is plenty of data, very little can be directly converted into action experience. What robots truly need is not just imitation of actions, but also information on how actions affect objects and how the environment responds.

To address this industry pain point, Dayao Robotics recently partnered with Nanyang Technological University's S-Lab to release ACE-Data-0, a multimodal embodied dataset, along with a paper titled 'ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine.' The report states that this upcoming dataset includes over 150 hours of home activity recordings, 17 million video frames, 200 task categories, 50 participants, 2 real-home environments, and more than 75,000 interaction clips.

More importantly, its core design shifts the organization of embodied data from videos that only show the appearance of actions to simultaneous recordings of vision, motion, sound, touch, object states, and task context within the same time and space. In other words, models need to learn why an action succeeds, when it fails, and how the physical world changes.

However, after reading the paper, the most interesting questions are just beginning to emerge:

  1. According to Daxiao Robotics, ACE-Data-0 belongs to a Level 5 dataset. How exactly are Levels 1 through 5 classified?

  2. ACE-Data-0 records natural human behaviors, but the human body, tactile sensations, and action spaces differ from those of robots. How can this data be transformed into capabilities that robots can truly execute?

  3. As human demonstrations, real-robot data, and simulated generated data all expand simultaneously, how will the three routes divide their roles? What is the next bottleneck for Embodied AI: is it still in the model, or is it shifting toward data?

Therefore, following the paper's release, 42nd Radio spoke with a representative from Daxiao Robotics about the multimodal embodied dataset ACE-Data-0, aiming to gain a deeper understanding of Daxiao's perspectives on embodied data, physical supervision, and industry evolution beyond the scope of the paper.

He judged that the industry bottleneck in the next stage is likely to shift from a pure competition of model architectures to whether models can obtain real, complete, and reusable physical supervision: "After different teams adopt similar architectures, performance differences are increasingly determined by contact, state, failure, long-range context, and cross-embodiment coverage in the data."

Where Does the High Information Density Come From

When asked why multiple modalities must be placed within the same spatio-temporal framework, the person in charge summarized the core issue of existing embodied data in one sentence: "The primary loss caused by fragmented data is not the absence of a certain modality, but the severing of physical relationships." While standalone video can identify action appearances, motion capture can reconstruct human movement, and tactile sensors can record contact, if these signals are not locked into the same timeline and spatial coordinate system, models cannot establish the complete chain of 'what was seen — how to act — what was touched — how objects changed.' They may excel at classification, retrieval, or short-term action imitation, but struggle to advance toward state prediction, failure localization, and long-horizon planning.

This is also the starting point for Dashiao's proposal of the "Embodied Model Information Density" classification. Dashiao Robotics provided a more specific explanation: L1 mainly consists of ordinary internet videos or images; L2 adds basic actions and language descriptions; L3 further includes structured states such as human body, hand, or object trajectories; L4 introduces multi-view, spatiotemporal synchronization, and robot control-related information; L5 requires continuous recording of vision, motion, sound, touch, object state, failure recovery, and long-range context in real open environments.

Based on this classification logic, ACE-Data-0 is defined as an L5-level dataset, based on data organization, information completeness, and physical causal density. The person in charge also emphasized that this classification does not claim that models have fully acquired generalization and reflection capabilities; these abilities still need to be continuously verified through subsequent downstream tasks, cross-environment testing, and long-term real-machine applications.

The core of this judgment is not complex: for robot learning, 1,000 clips of almost identical successful grasps may not be more valuable than one clip that completely records "approach-contact-slide-adjust-steady grasp." The truly scarce signals often appear at the moment when action results change.

Therefore, ACE-Data-0 did not require participants to strictly follow scripts step by step, but only gave goal-level instructions. For example, 'prepare a cup of tea and deliver it to the dining table.' Participants can choose cups, routes, and sub-task sequences on their own, and naturally encounter processes such as waiting, hesitation, returning to the previous step, and handling unexpected situations. For the same task, some people boil water first, while others look for tea bags first; some complete the grasp in one go, while others readjust after objects slide.

A person in charge at Dashiao Robotics explained the value of such processes using pouring water as an example: When people find that the rim of the cup is not aligned, they will often pause first, reduce the angle of the kettle tilt, readjust the position, and then continue pouring. A standard successful trajectory can only tell the model "what should be done under ideal conditions," while the complete process also provides premonitions of failure, pause conditions, cause judgment, and recovery strategies. This gives the model the opportunity to learn what signals mean that an action is about to lose control, when adjustments are needed, and which steps need to be retreated to for re-execution.

The value of failure data is not to teach robots to make mistakes, but to enable them to learn to recognize errors, locate errors, and correct them autonomously," said the person in charge. For home environments where it is difficult to always maintain a standard initial state, recovery capability is often more important than replicating a successful trajectory.

These differences are easily treated as noise in highly standardized data collection, but for robots entering real homes, they constitute reality. The difficulty of a long-range task is not just that the video is longer, but that past actions continuously change future options: where the cup was placed, whether the countertop has been cleared, which step has been completed, and whether an object left the field of view and then reappeared. Models need to remember goals, track states, and continue acting when plans are interrupted.

To support such natural behaviors, ACE-Data-0 designs three categories of tasks. Atomic-level human-object interaction segments are approximately 3 minutes each, covering operations such as pouring water, washing, wiping, cutting, folding, and organizing; long temporal activity chains last about 20 to 30 minutes, placing subtasks like fetching items, washing, chopping, cooking, plating, and tidying on a continuous timeline; human-scene interaction segments are approximately 5 minutes each, recording actions such as walking, sitting, lying down, exercising, operating furniture switches, and transitioning postures. Even the shortest data collection is measured in minutes, rather than the few-second clips common in traditional hand-object interaction datasets.

Turning Homes into Data Factories

To record the complete process, having a large number of cameras is just the first step. The real challenge lies in determining whether what different devices observe and measure truly corresponds to events occurring at the same moment and location.

Fine-hand manipulation and room-scale activities impose nearly opposite requirements on sensors. Fingertip contact demands close-range, high-density observation; movement between a kitchen, living room, and bedroom requires broad coverage and stable global coordinates. To address this, Dashiao designed its ambient data collection engine, ACE, with two complementary configurations.

The desktop-scale system covers a workspace of approximately 30 square meters. Eight close-range GoPro cameras observe the operation area from different angles, while sixteen OptiTrack cameras handle optical motion capture. Tactile gloves record pressure changes in the palm and finger regions. It focuses on how grasps form, when fingers slip, and how tools and objects undergo subtle rotations.

The room-scale system covers a complete home environment of approximately 200 square meters, including the kitchen, dining room, living room, and bedroom. Eight fixed RGB cameras and twelve OptiTrack cameras cover the activity space, ensuring that any point within the activity area can be observed from at least four external viewpoints, even in the presence of furniture occlusion. Participants also wear first-person devices equipped with four fisheye cameras and IMUs, motion capture suits with 41 markers, and hand and haptic devices. External viewpoints supplement full-body and environmental information, while first-person perspectives retain local views that closely resemble robotic observation methods.

The two systems ultimately output not several unrelated files, but a unified multimodal timeline. Cameras, motion capture, tactile sensors, and audio devices each have their own clocks and sampling rates; a deviation of just a few milliseconds could result in the "hand in the video not yet touching the cup" and "tactile pressure already generated" occurring simultaneously. ACE uses the 60Hz motion capture clock as its time reference, estimating offset and drift by having cameras photograph QR codes displaying the host machine's clock. The paper reports that independent verification found camera time deviations to be less than 4 milliseconds.

This is not merely an engineering metric. During moments of contact such as grasping, colliding, sliding, pressing buttons, cutting, and rapid placement, a deviation of just tens of milliseconds can obscure the entire critical event. A representative from Daxiao Robotics stated that multi-device synchronization determines whether physical supervision is reliable: "The model might therefore mistakenly learn 'approach' as 'contact,' bind 'applying force' to incorrect states, or even learn reversed causal sequences."

Spatially, the fixed third-person camera, the constantly moving head-mounted device, the human body, hands, and objects are registered to the same world coordinate system. The median reprojection error of external cameras is less than 3 pixels, while that of head-mounted cameras is approximately 2 pixels. This allows researchers to locate corresponding human poses, object six-degree-of-freedom poses, contact pressures, and sounds from any video frame, as well as project three-dimensional points at the same moment into different camera views.

This engineering effort is not trivial. The paper reveals that one hour of data collection generates approximately 1 TB of raw data. The dataset provides four-channel first-person fisheye video, eight-channel third-person video, a forty-one-joint human skeleton, hand pose, SMPL-X parameters, object meshes and sixty Hertz six-degree-of-freedom trajectories, multi-source audio, full-palm pressure maps, as well as camera intrinsics, extrinsics, and synchronization tables. In addition to natural language descriptions, most annotations come from calibrated physical measurements rather than post-hoc model estimation on the video; text descriptions are generated by Gemini and then manually reviewed and corrected.

This is the most substantial change of ACE-Data-0 compared to ordinary event videos: vision is no longer the sole truth. Vision is responsible for presenting external changes, motion data explains how the human body and objects move, audio records collisions and equipment status, and touch directly informs the model whether contact has occurred and where the pressure distribution lies. They do not independently describe similar actions, but jointly record the same action within a unified spatiotemporal framework.

In Daxiao's vision, different modalities will eventually form a clear division of labor: vision is responsible for global understanding and action prediction, while tactile sensing provides closed-loop correction during critical contact phases. Grasping ordinary rigid objects can rely more on vision; flexible packaging, fabric folding, cable management, fragile objects, plug-and-pull assembly, and precise force control tasks require real tactile feedback. The significance of multimodal systems is not simply to stack sensors, but to let each signal play its role in the环节 where it excels.

30 Different Methods of Capability Verification

The ACE-Data-0 paper first conducts a more fundamental diagnostic: decomposing Embodied AI perception into three layers—low-level signals, scene composition, and interaction processes. It evaluates over 30 existing methods on reserved test data spanning 10 hours, observing which capabilities different models possess in real-home interactions and identifying directions for further improvement.

The first layer involves predicting tactile feedback from vision. Given only first-person videos, the model must determine when contact occurs, where pressure is distributed on the palm, and the magnitude of that pressure. Among three baselines, TouchAnything achieved a frame-level contact state accuracy of 0.7095, significantly higher than the other two methods; however, its contact region IoU was only 0.1646, and the volume IoU considering pressure magnitude was merely 0.1357. In other words, while the model is reasonably adept at judging 'whether contact occurred,' it has far to go in accurately answering 'where the contact happened and with what force.' This underscores that vision and touch are likely complementary rather than mutually exclusive: vision can predict contact trends, while touch provides direct evidence for pressure, friction, local slippage, and impending instability.

The second layer focuses on human motion recovery. The paper evaluated various methods including single-frame, temporal, scene-aware, multi-view, and first-person approaches. Results showed that some models could recover relatively accurate joint poses from single frames, but errors increased significantly when these poses were linked into long-term motion trajectories within a room coordinate system. Scene information primarily helped models determine a person's location within a room but did not necessarily improve local joint pose estimation simultaneously. Multi-view methods did not naturally outperform the strongest single-view method in this experiment, which the paper attributes partly to the limited availability of multi-view baselines, rather than dismissing the value of multi-view approaches entirely.

The third layer concerns hand motion during hand-object interaction. First-person views are closer to the hands but affected by head self-motion, fisheye distortion, truncation, and blur; third-person camera coordinates are stable, but hands appear smaller and are more easily occluded by the body or objects. On paired data, the best-reported world-coordinate trajectory error for third-person methods was 63 millimeters, whereas two first-person world-coordinate methods reported errors of 98.2 and 102.1 millimeters respectively. Although this does not constitute a strict ablation study of view angles for the same model, the results still point to a common issue: local finger pose estimation can be quite good, but the true differentiator is the ability to maintain a stable global trajectory over long sequences.

'Larger pixels do not equate to clearer spatial relationships,' said a representative from Daxiao Robotics. A first-person perspective is a continuously moving local window; the model needs to decouple camera motion, human motion, and object motion while maintaining a unified 3D world coordinate system. Fixed third-person perspectives are farther away but can continuously observe the overall relationship between humans, both hands, objects, and the environment. ACE-Data-0 records both types of perspectives simultaneously, providing exactly one set of physical ground truth for studying this complementary relationship.

These results make the value of the dataset concrete. A single robot task failure may not be due to the model being 'not large enough', but rather a lack of contact perception, an inability to restore object states, loss of global position, or accumulated drift over long time scales. Only by recording these intermediate signals can researchers move beyond final success or failure to ask further: at which layer did the error actually occur?

From Dataset to Data Infrastructure

The significance of ACE-Data-0 lies not only in providing new benchmarks for tactile inference, human motion, and hand motion recovery. First-person and third-person videos, human and object trajectories, audio, tactile data, and language descriptions are all within a unified spatiotemporal framework. This data can further feed into imitation learning, world models, and Vision-Language-Action (VLA) models, offering more complete supervision for transferring human demonstrations to different robot bodies.

Bridging the gap from natural human behavior to robotic execution requires solving the mapping between task intent, physical constraints, and specific embodiments. A person in charge at Daxiao Robotics outlined this path in three layers: first, extracting task intent and physical constraints from human data; second, mapping them to specific robot embodiments; and finally, correcting control errors through real-machine feedback.

"Human data is better suited for teaching 'what to do and why,' while robot data handles 'how exactly this machine does it.'" The person in charge stated. Human demonstrations provide task decomposition, object functionality, action consequences, long-range sequence adjustments, and failure recovery; real-machine data handles grasp force, joint control, end-effector trajectories, friction compensation, collision boundaries, and embodiment dynamics. By establishing general cognitive and behavioral priors with human data, and then completing embodiment adaptation and control alignment with robot data, the cost of collecting data from scratch for every robot and every task can be reduced.

From a broader industry perspective, the three main routes for Embodied AI data will complement each other. Natural human demonstrations are suitable for scaling up the learning of task intent, behavioral diversity, long-range structures, and common sense; simulation and world model-generated data can cover extreme cases, failure branches, and numerous environmental changes at low cost; real robot data provides embodiment dynamics, contact control, execution errors, and safety boundaries. The person in charge judged that generated data cannot fully replace real physical interactions, and the task structures, open-ended variations, and recovery strategies contained in natural human data still hold significant value that has not yet been fully utilized.

This also implies that the next round of competition in Embodied AI may no longer just be about competing on model architectures. As different teams adopt similar VLA or world model architectures, performance differences will increasingly depend on whether the data retains information on contact, state, failure, long-range context, and cross-embodiment details. The industry's method of measuring data will gradually shift from "how many hours and trajectories were collected" to whether this data can be synchronized, converted, verified, and reused.

Rather than building an all-encompassing mega-dataset, Daxiao hopes to promote a set of interfaces that allow multi-source data to understand each other. Different robots, sensors, control frequencies, coordinate systems, and action definitions require unified time synchronization, spatial coordinates, human-to-robot action representations, object states, quality evaluation, and language descriptions to avoid forming new data silos as scale increases.

"What will truly become the industry interface may not necessarily be the team with the most data, but rather open standards and toolchains that enable low-cost access, conversion, validation, and training usage of multi-source data," a representative from Daxiao Robotics stated.

From this perspective, the '0' in ACE-Data-0 is more like a starting point. It transforms a real home into a collection environment capable of sustainable recording, unified calibration, and precise synchronization, while translating human perceptual, tactile, and motor experiences from daily life into physical knowledge that robots can read. When these experiences can flow across scenarios, tasks, and embodiments, high-quality data can evolve from a resource owned by individual teams into shared infrastructure driving the development of world models, VLA (Vision-Language-Action) models, and general physical intelligence.