Unitree Open-Sources 'Trump Card' Technology, Sending Shockwaves Through the Robot Industry
On the evening of September 15, Unitree announced the open-sourcing of the UnifoLM-WMA-0 architecture, a world model-action architecture under the UnifoLM series specifically designed for general robot learning. This open-source initiative provides a solution for cross-type general robot body learning and is expected to advance the development of the global Embodied AI industry.
The core component of the UnifoLM-WMA-0 architecture lies in a world model that can understand the physical laws governing the interaction between robots and their environment.
According to reports, this world model possesses two core functions:
- Simulation Engine, which operates as an interactive simulator, providing synthetic data for robot learning;
- Policy Enhancement, which can be connected with an action head to further optimize decision-making performance by predicting future interactions with the physical world.
In terms of training strategy, Unitree Technology performed specialized fine-tuning on video generation models using the Open-X dataset to adapt their generation capabilities to robotic work scenarios. The model generates corresponding future action videos based on images and text instructions. The generation results of the fine-tuned model on the test set are as follows:
Furthermore, under the UnifoLM-WMA-0 architecture, the world model is not limited to a single mode; it supports two operation modes—decision mode and simulation mode—providing a certain degree of flexibility.
In the decision mode, the model can predict information about future physical interactions between the robot and its environment to assist policy generation. This enables the robot to "think" ahead about the potential consequences of actions when performing tasks, thereby making better decisions.
In the simulation mode, the model can generate high-fidelity environmental feedback based on the robot's actions. It effectively simulates an interaction environment for the robot that is very close to reality.
The complete system architecture and workflow are as follows:
Furthermore, the team completed the model training based on five open-source datasets from Unitree Robotics.
Test results show that, as a simulation engine, the model can achieve controllable interactive generation based on a "current image" and a certain number of "future robot actions".
A comparison of the generated results with the original video is shown below:
More importantly, the model possesses the ability to continuously interact and generate for long-horizon tasks. This means it can not only handle immediate tasks but also plan and execute complex tasks that require multiple steps, significantly enhancing the practicality of robots in real-world scenarios.
For instance, in a task involving placing a black camera into a packaging box, based on the world model's prediction of future action videos, the robot first determines the orientation of the camera, places it into the recess of the box, and finally closes the lid in a specific direction. This demonstrates the model's real-time prediction capability during environmental interaction.
In an item organization task, the robot first identifies the scattered items on the desk—specifically an eraser and a pen in actual scenarios. It then distinguishes where each item should be placed within the box layout: the eraser goes into a small compartment, while the pen is placed in a larger one. After positioning both items, the robot closes the box.
During the process of stacking wooden blocks, it picks up the blocks in red-yellow-green order and adjusts their angles when placing them to ensure all three blocks are aligned.
According to Unitree's official website, UnifoLM has previously been integrated into the Unitree G1. Regarding the complete open-source release of UnifoLM-WMA-0, Unitree Technology stated that it will continue to update in the future.
For current robots, models are a significant barrier to their entry into household life. Previously, Wang Xingxing also stated at the Bund Summit that current robot hardware is entirely sufficient; the biggest issue lies with the models themselves. The inherent capabilities of the models are insufficient, preventing them from effectively utilizing the hardware.
Furthermore, at the World Robot Conference held in August, Wang Xingxing mentioned that controlling robots to mimic execution routes using pre-trained robot action videos might develop faster and have a higher convergence probability than VLA (Vision-Language-Action) models.
Through the open-source release of UnifoLM-WMA-0, developers across the industry can further optimize robot control algorithms based on this architecture.
This also gives Unitree the opportunity to occupy a more important position in the next stage of robot development.
