IROS On-Site Dialogue with the ULTRA Team: What's Still Missing for Robots' 'Movement Freedom'?

Dexterous hands, drive motors, control algorithms, and large models—the components forming the "body" of embodied AI—are becoming increasingly precise.

Consider moving boxes. A robot's dexterous hand must locate contact points, its feet brace the body, and motors drive it to squat while the torso adjusts posture in response to shifts in center of gravity. It then carries the box, coordinates balance, and moves slowly. For a human familiar with this labor, these movements require only a few weeks to develop coordination and muscle memory, without needing to think through each step individually.

Humanoid robots are learning similar capabilities. Researchers provide motion demonstrations, and the resulting "motion tracking" allows robots to follow complex movements. However, they often require continuous external reference to guide their next actions.

But limited motion demonstrations are far from enough. To achieve human-like control gaits and physical interaction in the real world, autonomous whole-body locomotion is an inevitable goal.

So, how do we transition from control to general action capabilities? How can we explain "muscle memory" and "autonomous coordination" to robots? And if such memories exist, how are they written into a robot's body?

ULTRA may offer a solution. During IROS 2026, He Xia Lin and Xu Si Rui's team, with their paper ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation, stood out among numerous submissions and became one of the teams shortlisted for the Best Mobile Manipulation Paper Award at IROS 2026.

He Xia Lin is a third-year Ph.D. student in Computer Science at UIUC (University of Illinois Urbana-Champaign), advised by Professor Saurabh Gupta. He completed his undergraduate studies in the ACM Honors Class at Shang Hai Jiao Tong Da Xue, where he collaborated with the SJTU APEX lab on applying reinforcement learning to quadruped robot locomotion under the guidance of Professor Zhang Wei Nan. As a senior, he served as a research intern at UCSD, mentored by Professor Xiaolong Wang and collaborating closely with Professor David Held. His research focuses on reinforcement learning, robotics learning, computer vision, and control theory.

Xu Si Rui and He Xia Lin are both Ph.D. candidates in Computer Science at the University of Illinois Urbana-Champaign (UIUC), advised by Professors Yuxiong Wang and Liangyan Gui. They completed their master’s degrees at UIUC and their undergraduate studies at Peking University, with a research focus on building interactive embodied agents capable of understanding, practicing, and mastering physical interactions like humans.

Under the guidance of Professor Liangyan Gui in the UIUC Department of Computer Science, He Xia Lin and Xu Si Rui collaborated on the development of ULTRA. This research aims to enable robots to learn whole-body coordination from human demonstrations and then act based on goals and their observations of the environment.

During a live forum presentation at the conference on the 28th, He Xia Lin and Xu Si Rui noted that current embodied AI must overcome two major hurdles: continuity in task execution and generalization across diverse scenarios. When given only a goal, a robot must autonomously organize intermediate actions. If the command style changes or the environmental observation shifts, the skills it has already learned must be able to generalize effectively to out-of-distribution scenarios.

In an interview, Xu Si Rui described this process as compressing and retaining experiential knowledge from human data so that different tasks can call upon learned skills. However, he pointed out that connecting these skills to a robot’s perception of a complex world remains an area requiring further exploration.

Robots Forming "Muscle Memory"

Achieving autonomous, multi-functional loco-manipulation remains a core obstacle to making humanoid robots truly practical. The key lies in creating a formal logic for embodied AI to develop a form of "muscle memory." However, limitations exist: on one hand, data supporting action retargeting for robots is often scarce or of low quality, making it difficult to scale to large skill libraries. On the other hand, such methods rely on tracking predefined motion references rather than generating behaviors based on perception and high-level task specifications.

Human demonstrations contain rich operational experience, but this experience first belongs to the human body. For a robot with shorter arms or different joint ranges of motion, the same posture might not allow it to reach a box. Adjusting the same movement sequence to fit a robot's specific body is known as "action retargeting."

"We need to transform so-called human motion data into a state that robots can use."

ULTRA proposes a unified framework, comprising two key components: a physics-driven neural retargeting algorithm and a multimodal controller. The neural retargeting algorithm transfers large-scale motion capture data to the robot’s body while maintaining physical plausibility during complex contact interactions. The unified multimodal controller supports both dense reference inputs and sparse task specifications, with perception conditions extended to include ego-centric data inputs.

This approach allows ULTRA to distill a general-purpose motion tracking policy into the controller, compressing motor skills into a compact latent space and leveraging reinforcement learning to enhance the robot’s generalization capabilities.

The challenge lies in the fact that even when poses are closely matched, actions can still fail. For instance, if a human demonstrator’s hand is in contact with a box, the robot might leave a gap; similarly, foot placement may appear correct but result in sliding along the ground during execution. Since moving a box requires continuous hand-object contact, these discrepancies directly impact task completion.

What are the deeper underlying issues?

Although "foundation model-style" control has enabled the distillation of large-scale motion data into reusable expertise, scaling capabilities solely through massive, heterogeneous mobile manipulation datasets has proven ineffective. Furthermore, offline distillation is constrained by the state coverage of the teacher model’s trajectory rollout, causing problems to become particularly pronounced in high-dimensional robot–object interactions.

To address this, ULTRA places action conversion within a virtual training ground governed by physical laws—specifically, a physics simulation environment. The robot repeatedly attempts tasks within this simulation, where the training system evaluates performance based on body movements, box positions, and hand contact states, guiding the model to improve incrementally. This trial-and-error learning method constitutes reinforcement learning.

During this process, actions are subjected to physical calculations within the simulation. Whether a hand touches a box or feet make contact with the ground directly influences subsequent events.

The team also allows robots to adjust certain postures, prioritizing critical areas such as hands and feet. When the robot's body configuration changes, its stance and reaching methods may need adjustment to lift the same box.

Xu Si Rui emphasized that learning in physics simulations aims to reduce issues like inadequate contact and visual clipping between the virtual body and objects. Resolving these problems ensures that the resulting actions are suitable for the next stage of training.

"Human data is based on human anatomy, which differs from robotic structures; therefore, human motion data must be converted into states usable by robots. Our approach has two key features: first, using neural networks to perform this mapping; second, training this network within a physics simulation environment to achieve motion retargeting. Through the simulation, we aim to ensure that the retargeted motions are physically feasible, avoiding issues such as object clipping and insufficient contact."

Comparisons in the study also focused on these challenges. On both large boxes and suitcases, ULTRA-generated motions demonstrated greater robustness compared to other methods, exhibiting less foot sliding and fewer instances of hands losing contact with objects during transport.

This method can also augment training data. By scaling boxes larger or smaller and adjusting the spatial amplitude of existing motions, already trained models can still attempt to convert demonstrations into new actions without requiring retraining from scratch. These operations expanded the dataset used in the paper to approximately six times its original size. A single demonstration of lifting a box thus generates more practice scenarios, enabling robots to interact with objects of varying sizes and different body positions.

Among Models, There Is Always a Mentor

With demonstration data available for robots, how can it be transformed into a robot's skill?

ULTRA first trains a "teacher" model, and then a "student" model. The terms "teacher" and "student" here define the relationship between the models. Although both are control models that determine how the robot acts next, their most important difference lies in the information they receive and their respective goals. Xu Si Rui believes that the "teacher" model has more comprehensive information in simulation, understands the states of the robot and objects better than the "student" model, and can see detailed action references, allowing it to focus on following demonstrations. The "student," on the other hand, must organize actions suitable for the current task with less prompting and incomplete observations.

How do you pass on a teacher's expertise to students? "Distillation."

The team trained a teacher model to imitate redirected human motions. During training, the team randomized physical parameters in the simulation environment and applied external forces to the robot, enabling the teacher model to learn how to recover from disturbances. When complete motion references were unavailable, the team distilled these skills into a shared representation state known as the "motion latent space".

This "motion latent space" can be understood as a set of compressed, more compact movement experiences. Whether the input is a detailed demonstration or a target destination for the box, the model will combine this experience with the robot's proprioceptive state to autonomously determine its next motion strategy. If given different commands, the robot can still leverage these previously learned experiences, much like reusable "muscle memory."

During this period, the student policy receives the robot's body state information and history, along with some optional inputs. For example, short-term reference poses for action tracking, long-term object targets, point clouds generated by an onboard depth camera, or object states provided by a motion capture system. These reference inputs collectively support better action performance for the robot in in-distribution scenarios.

To ensure the student model demonstrates generalization in out-of-distribution scenarios without these references, the team designs random masking of partial signal inputs during training. This forces the policy to practice executing tasks with incomplete information, creating a form of 'masking.' Sometimes, the robot is not given complete motion references or certain object details, requiring the student model to use remaining clues to complete actions and continuously refine its performance. During this process, the teacher model provides guidance to help it retain learned capabilities. The team also alters physical conditions in the simulation or pushes the robot, allowing the teacher model to practice recovery after interference. The experience accumulated this way includes both how to follow motions and how to correct deviations.

Even so, the information provided to robots in the real world is often less comprehensive. An operator might only specify the target position of a box, without providing intermediate steps. The robot's understanding of objects may also rely solely on depth cameras carried by the robot itself, which can measure distance.

The concept of 'unified multimodality' becomes concrete here.

ULTRA has designed a novel multimodal controller, which He Xia Lin and Xu Si Rui summarize along two dimensions: perception and control.

From a perception perspective, one approach is observing the world through cameras mounted on the robot body; another is using additional data acquisition equipment to obtain more precise motion information. From a task and control perspective, one approach is providing full-body demonstrations for the robot to follow; another is providing only sparse conditions and observations, requiring the robot to perform full-body manipulation. Perception methods and control methods can be combined in pairs to form different usage patterns.

On one hand, ULTRA tells the robot what to do by specifying detailed body movements, object positions, or the robot's target location and orientation. On the other hand, external motion capture equipment can inform ULTRA of object positions, while the robot can also use its own cameras to observe, enabling a better understanding of its environment. These methods can be combined and processed by the same model. When detailed guidance is present, the robot follows it; when only a destination is given, the robot must fill in the intermediate actions itself.

In the demonstration, you simply click the box you want to move, like playing an electronic game, and then select its target position.

Training ULTRA to Surpass Its Teacher

The teacher’s experience supports the student’s long-term progress, but imitation alone is not enough. An excellent student should surpass the original through practice. The same applies to robots.

A demonstration record provided by the teacher model represents past experience—a process moving from a starting point to an endpoint. Specifically, in a handling scenario, if the box’s position changes relative to the demonstration record, or if the destination no longer matches the memory, the student model should not be at a loss even in such out-of-distribution scenarios.

How? Training.

Just as students receive extra exercises after learning from a teacher, the ULTRA team alters the starting points and targets. When the robot makes progress toward the target or completes the task, it receives a reward. Additionally, the student continues to receive some guidance from the teacher, reviewing old knowledge to gain new insights, thereby avoiding forgetting old skills while learning new situations.

These exercises also fully account for obstacles such as lighting conditions and occlusions, including scenarios where visibility is poor. The team introduces errors, missing data, and occlusions into simulated visual information, allowing the model to adapt to less-than-ideal observation conditions. These improvements are particularly evident in tests that deviate from the original training objectives.

The team tested the ULTRA model, trained in one simulation environment, in another. Faced with new targets involving random offsets, when using simulated onboard visual information, the number of completed tests increased from 5 to 9 out of 20 after additional training; when object positions were provided directly, successful attempts increased from 4 to 12.

The results show that ULTRA outperforms baselines across nearly all metrics and categories, while also mitigating mesh-intersection issues. The team attributes this to the effectiveness of physical retargeting, which ensures stable foot placement and maintains hand-object contact during lifting. Furthermore, under continued training of the teacher model, ULTRA’s motion tracking success rate in out-of-distribution scenarios is significantly higher than that of versions trained on its original dataset, demonstrating substantially improved generalization compared to OmniRetarget.

These results support the team’s assessment. Distillation from the 'teacher' to the 'student' model enables performance transfer: ULTRA’s ability to learn corrective capabilities with contact awareness using privileged states and dense references indeed yields higher success rates. Without retraining, ULTRA’s retargeter can diversify motions and apply changes across the entire trajectory, producing temporally consistent variations.

However, experimental data indicates that ULTRA’s success rate remains unstable when facing new, out-of-distribution situations. At the same time, what are the costs of allowing the policy to receive more diverse and combined inputs? For example, does it lead to increased latency or other performance degradation?

At the IROS conference, an audience member raised this question to He Xia Lin. He noted that in motion tracking tests within the training scope, the unified model’s performance declined. Nevertheless, different control modes did not appear to impair the model’s generalization capabilities for out-of-distribution cases.

'We believe that once a shared latent space is learned, the policy can maintain appropriate behavior across different input combinations. This is one of the key ideas of the paper.'

Perception Remains the Hardest Challenge

The team conducted deployment tests on the Yu Shu G1 robot, experimenting with various control methods. When provided with complete reference data, the robot follows the reference; when the reference is removed, it acts autonomously based on the object’s target position, and operators can also adjust the target step-by-step via keyboard. During this process, ULTRA can utilize external motion-capture equipment or rely on its own depth camera to perceive object information during control.

Based on the project demonstration, ULTRA can perform tasks such as lifting boxes with both hands, transporting boxes and suitcases, and kicking objects. The motion skills acquired through training remain usable across different targets and command methods. Using fine-grained keyboard control combined with motion capture, operators can gradually adjust target positions via directional keys, while the controller maintains overall movement coordination.

During conference presentations, the team highlighted the impact of perception conditions.

In trials where the robot followed complete motion references, the task success rate was approximately 73%, achieving 19 successes out of 26 tests. When motion references were removed and object states were provided by external devices for vertical and lateral movement tests, the robot succeeded in eight out of ten attempts for the former and nine out of ten for the latter. Relying solely on its own camera observations reduced the success rate to 50–60%. These results confirm that even without frame-by-frame guidance, the robot can execute actions under ULTRA’s command based on target instructions. However, reliable environmental perception significantly affects task execution outcomes.

The most challenging aspect of transitioning from simulation to real-robot deployment remains perception and control interfaces.

As the robot’s eyes and cerebellum, camera observation and body control operate at different rhythms. Visual systems require time to process images, while bipedal robots must continuously update movements to maintain balance. The body cannot wait for the next image processing cycle before deciding how to stabilize itself.

Consequently, the team implemented filtering and interpolation techniques. These operations reduce jitter and noise in observations and synchronize information with varying update frequencies. Additionally, the robot must handle scenarios where arms or objects obstruct the camera view. These seemingly minor processing steps determine whether learned motions can be executed stably on physical hardware.

Xu Si Rui recalled that in the early stages of the project, there were many cases where simulations succeeded but real-world attempts failed. As R&D progressed, the team worked to make the control models more stable while also improving the coordination between perception and control on physical robots. He was primarily responsible for algorithm and model development, while his co-authors focused more on deploying the system on actual hardware.

"What we can currently achieve is compressing collected human data into a reusable latent space, and then using this latent space to construct and reuse various skills for general manipulation. However, this raises another issue: how do we connect these skills with observations of the world? For example, observing that there is a cup here and then reaching out to grab it remains an area worth exploring. In complex environments, we need to invoke different modules for object detection, segmentation, and recognition, and then feed these results into action interfaces to call appropriate actions. I believe we are still in a relatively early stage."

The reasons for failures in certain scenarios are also quite specific. Real-world friction conditions differ from simulations, which may cause boxes to slip from the robot's grasp; inaccurate camera measurements or occluded lines of sight can lead the robot to misjudge objects; and excessive external disturbances may exceed its trained recovery capabilities. In short, there are indeed areas where performance falls short of expectations.

"The area that has not met expectations is its lack of certain dexterous manipulation skills, such as grasping more complex geometries with a dexterous hand. Currently, it mainly handles boxes or objects with shapes similar to suitcases, meaning the variety of operable objects remains relatively limited. The primary reason is that the tasks lack sufficiently precise data for the entire method to utilize; from the perspective of the complete methodology, there is also a lack of more accurate observation. At present, we use a relatively simple PointNet to encode point clouds generated from depth images. However, dexterous manipulation may involve more occlusions and more complex spatial relationships. We may need a stronger model to provide corresponding information before connecting to the motion control interface."

In fact, whether an object's position is reliable, whether the hand can maintain contact, and whether the body can remain stable after being disturbed all affect whether the next task can be completed. These issues pull "autonomy" back into every observation and action update, awaiting future resolution by the team.

**Constructing the Future "Cerebellum"

If we continue forward, what does ULTRA most need to improve?

Xu Si Rui pointed out the limitations of dexterous manipulation. He acknowledged that the types of objects currently operable by the ULTRA system remain relatively limited, primarily consisting of boxes and luggage items approaching box-like shapes. The ability to grasp more complex geometries with a dexterous hand has not yet met the team's expectations.

In his view, this involves both data and what the robot can perceive. Finer tasks require more precise demonstrations for the model to learn; complex operations with greater occlusion and more intricate spatial relationships demand higher visual capabilities.

Currently, ULTRA converts depth camera data into point clouds to describe object shape and position. The processing selects the area in front of the robot, excludes the ground, and identifies the primary cluster of points as a bounding box. This approach supports current experiments but requires stronger recognition capabilities to operate in complex environments filled with diverse items.

For example, if you see a cup, you reach for it. The robot must first locate the cup among other objects, determine its position, and then pass this information to the motion control model to execute appropriate actions.

Xu Si Rui believes there is still much to explore in this process.

Robots could also benefit from additional sensory input, such as the force applied by the hand and how to adjust grip and movement upon contact. Although the research has not yet incorporated force feedback or compliant control, He Xia Lin suggests that integrating these capabilities is a promising direction for future work.

This work, which spans data, algorithms, and deployment, is a process of accumulated effort. Xu Si Rui is currently pursuing a Ph.D. at UIUC, with research covering computer vision, machine learning, and robotics. He notes that he has worked in this field for several years, and many aspects of ULTRA reflect the continuity of his previous work.

"This work carries many traces of my prior research."

When the conversation turned to entrepreneurship and career choices, he mentioned that paths such as starting a company or pursuing an academic position were all under consideration. However, he noted that a single paper does not necessarily translate into commercial value; instead, he hopes for the team to sustain output in one direction. Demonstrating incremental improvements over time is what truly showcases the team’s capabilities.

Regarding its future positioning, Xu Si Rui stated that providing robots with an advanced 'cerebellum' is crucial. In his vision, ULTRA can connect to more powerful visual and task-planning systems: the upper-level system understands tasks and sets motion goals, while the lower-level motion module handles physical coordination during specific actions. This means that simply lifting a box is just one step. Organizing learned actions sequentially based on task requirements and adjusting them amid environmental changes requires better integration between the upper-level system and motion capabilities. The team plans to integrate Agent or VLM models so that existing skills can serve longer, more complex tasks.

"We aim to provide a cerebellum, or a low-level motion module, combined with higher-level deliberative thinking and systemic capabilities to accomplish task planning over longer time horizons. ULTRA establishes an interface that can accommodate various perception modules and control methods. Similarly, we can attempt to integrate stronger Agent or VLA modules to enable longer tasks, organizing different skills temporally. Through this interface, we hope this method can be applied in more complex scenarios."

Returning to the initial box-lifting example, ULTRA demonstrates that the practical potential of embodied applications is approaching. Even without action references, the coordination skills the robot has learned continue to function. The blueprint for full-body autonomous mobility depicted by ULTRA may be closer than we think.