Tsinghua and Peking University Teams Release SkelWAM for Zero-Shot Robotic Arm Transfer
SkelWAM enables an operation policy trained on one robot to be directly deployed onto other robots with different structures. The research team designed a common geometric representation based on a unified arm skeleton. The model receives visual input, reads proprioceptive states, and predicts subsequent actions, all based on this representation. The policy was trained exclusively on Franka; when transferred to ten previously unseen robots, the model weights required no modification whatsoever.
On September 18, the School of Advanced Manufacturing and Robotics at Peking University, Qing Yu Ke Ji, and the Institute for Interdisciplinary Information Sciences at Tsinghua University released the paper "SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation." The paper has five authors. Pengjun Niu and Yujia Xie are co-first authors. Zhao Xing, Assistant Professor at the Institute for Interdisciplinary Information Sciences at Tsinghua University and head of MARS Lab, and Liu Ke, Assistant Professor at the School of Advanced Manufacturing and Robotics at Peking University, are corresponding authors.
World action models bring the predictive capabilities of video models into action generation: the model first predicts future visual changes, then outputs actions based on those predictions. However, transferring such models across different robots encounters dual obstacles. On the vision side, new robots vary in appearance, size, and imaging geometry. On the action side, control coordinates, degrees of kinematic redundancy, and the characteristic that a single tool pose may correspond to multiple arm configurations also differ.
In the past, handling the visual side mostly involved painting the source robot's appearance into the target observation, or synthesizing changes in the robot's appearance and viewpoint during training; on the action side, outputs were typically unified into end-effector poses, followed by a single inverse kinematics redirection for the target robot. This approach can align tool poses but cannot control the full-body postures required to achieve those poses.
The team's answer is a 25-dimensional skeleton state, where visual observations, robot states, and action goals share the same representation. The entire training process uses only demonstration data from the source robot.
01 Control Interfaces Evolve Alongside the Body
The core challenge of cross-body transfer lies in the control interface. For the same task of moving a gripper to an identical spatial point, a 6-DOF robotic arm and a 7-DOF robotic arm output completely different joint angles; redundant-DOF arms can even keep the tool stationary while adjusting the overall posture of the arm. If switched to a cable-driven flexible arm like the Qing Yu Ke Ji A03, the control output becomes curvature parameters, which are entirely different from the joint angles used for rigid arms.
The team identifies two gaps arising from differences across robot bodies. On the visual side, appearance, body shape, and the geometry visible in the camera change with the body; the same cup produces two different images in the cameras of two different robots. On the action side, control coordinates, degrees of redundancy, and the full-body configurations that determine achievable tool poses also vary by body. With both gaps present simultaneously, addressing only one is ineffective.
Existing methods mostly address only one side. Mirage renders the source robot's body within the target scene. RoVi-Aug synthesizes diverse robot appearances and viewpoints during training via data augmentation; both focus on vision. On the action side, common approaches involve uniformly outputting end-effector poses for the target robot to solve via inverse kinematics, or learning a cross-embodiment action latent space. These can recognize visual targets to guide grippers to desired positions, but leave the overall arm posture to be solved independently by the target controller.
SkelWAM's approach is to have observation and action share the same skeletal representation. The pose of the tool end-effector and the posture of the entire arm are described by the same data stream. When migrating a model trained on a source robot to a new device, only the calibration parameters and kinematics solver are replaced. The target robot provides only geometric calibration, camera parameters, and a kinematics controller. Model training and policy learning are all completed on the source robot; the target side does not use any demonstration samples.

The team validated the approach on 10 target robots: nine rigid robotic arms covering six-degree-of-freedom, seven-degree-of-freedom, and lightweight models, plus one cable-driven flexible arm. The test followed LIBERO's benchmark for objects, language instructions, reset conditions, and success criteria, with the metric being the success rate of completing an entire task.
02 A Skeleton to Clarify Posture
The core of SkelWAM is a fixed-length vector, which the team calls skeleton state, totaling 25 dimensions. Researchers performed six sampling points along the arm's centerline, from the base to the tool center point (TCP), with the final sample located at the TCP. The system calculates the offsets of the first five sampling points relative to the last point and normalizes them using the sampled centerline length; these fifteen dimensions describe the arm configuration. The remaining ten dimensions carry task-related information: three dimensions represent the TCP position in the task coordinate system, six dimensions provide a continuous rotational representation of the tool orientation, and one dimension corresponds to the gripper command, where a positive value indicates opening and a negative value indicates closing.
The body offset is normalized based on the centerline length, decoupling arm posture from the total body length and global position. This allows both six-degree-of-freedom and seven-degree-of-freedom arms to be described using the same set of values. The two normalization logics are independent: body offsets are based on the sampled centerline length, while tool positions are mapped to the task coordinate system. A set of per-dimension normalization parameters is then fitted based on the source data. After training, these parameters are fixed and remain unchanged when deployed to the target robot.
The visual frames received by the model are also generated by this geometric structure. The team removes the pixels of the original robot from the source image, fills in the background, and then projects the skeleton back onto the image. An external viewpoint draws a centerline connecting six sampling points, enabling the model to perceive the trajectory of the arm within its workspace. A wrist-mounted viewpoint relies on a camera bound to the tool coordinate system to render the outline of parallel-jaw grippers with perspective projection, facilitating the model's observation of the geometry of objects near the tool and the state of the gripper. The external viewpoint retains global scene information, while the wrist-mounted viewpoint focuses on the area near the tool; both views are defined by the same skeleton, ensuring that visual information and action targets remain aligned along the timeline.

The policy output provides only skeletal-level motion intent, with actual execution handled by the solvers built into each target robot. The solver prioritizes ensuring that tool pose and gripper commands meet requirements, treating the skeleton centerline merely as a soft posture constraint. The objective function consists of four weighted offset terms, augmented by a temporal continuity term for adjacent steps, without strictly forcing every sampling point on the centerline to match. Rigid robotic arms employ constrained joint inverse kinematics for solving, while cable-driven flexible arms use piecewise constant curvature parameters. Before initiating solving in the simulation environment, scale alignment is performed: body offsets are rescaled according to the length of the target robot's sampled centerline, and then tool positions within the task coordinate system are added.
The state variables also incorporate interaction feedback. The simulation bridge initiates computation from measured geometric information without pre-assigned intentions. For rigid arms, the body intention from the previous frame is retained, and the next frame's state is updated based on the measured tool position, orientation, and gripper commands. For flexible arms, both the body configuration and interaction orientation intentions are retained, with only the measured tool position and gripper feedback being updated. When the solver computes feasible solutions, it uses the robot's real-time measured configuration. The entire control system employs a receding horizon scheme: after executing a motion segment, new observations are used to replan.
03 Two Experts on Video and Motion
The model body comprises two independent expert modules, each with thirty layers, responsible for video and actions respectively. Both are built upon diffusion Transformer modules, with weights initialized from the Wan2.2 video backbone. The 25-dimensional state encoder, action input/output projection layers, and gripper event head are newly initialized and trained solely using demonstration samples from the source robot.
Within each pairing module, action queries attend only to the key-value pairs of the current video frame and other action tokens; video tokens from future frames do not participate in action attention calculations. During training, the video branch provides prediction supervision: the video expert learns to infer subsequent visual changes, thereby integrating dynamic scene information into action learning. In the inference phase, this video branch is no longer involved in the generation process. At inference time, the model encodes the current observation only once, caching the key-value pairs of the current frame's video so that action experts can directly reuse the cached content during sampling.
The training is divided into two stages, using only source robot data throughout.
The first phase consists of 6,000 update rounds, jointly training the last four modules of the video expert, the video output head, and the action expert. The optimization objective is a weighted future frame prediction loss combined with a 25-dimensional skeleton action loss.
In the second phase, video expert parameters are frozen while the action pathway is trained separately for 24,000 update steps. Continuous 24-dimensional geometric quantities are optimized using flow matching. Gripper open/close actions are handled by a factorized event head, which predicts whether a state switch occurs at the current step and the timing of the first switch. By decoupling gripper events from continuous geometric variables, discrete semantics like opening and closing are prevented from being smoothed over.
During inference, the model outputs 32-step absolute skeleton states in one go via 20-step flow matching. Each action block allows at most one gripper switch; additional switches are deferred to the next planning round. After executing the first 10 steps, the system re-collects observations and generates a new action block, forming a closed-loop control cycle.
04 Results Across Ten Robots
Building on the original LIBERO task suite, the team developed a new cross-robot evaluation benchmark called LIBERO-Cross10. This benchmark retains the original scenes, objects, language instructions, reset rules, and success criteria, but replaces the robot body, gripper structure, mounting method, and camera configuration to create ten distinct test environments. It comprises ten fixed tasks covering four subsets: spatial interaction, object manipulation, point-to-point targeting, and long-horizon execution, with three tasks each for spatial and object categories, and two each for point-to-point and long-horizon categories.
The test utilizes ten diverse robots, covering mainstream industry models and novel structures. Nine are rigid manipulators across three categories: industrial arms (JAKA mini2, UR5e, UR10e), high-redundancy arms (Kinova Gen3, xArm7, Rizon4, IIWA), and lightweight arms (ARX-L5, ViperX300). The final unit is the Qing Yu Ke Ji A03 cable-driven soft arm, enabling comprehensive testing of both rigid and soft robotics. Experiments were conducted in MuJoCo and robosuite simulation environments. Each robot completed 100 test rounds, totaling 1,000 trials across all devices, ensuring sufficient and balanced data samples.
Model training relied solely on source device data, using only 467 successful demonstrations collected from the Franka manipulator across the ten tasks. The dataset was split by trial: 447 samples for training and 20 for validation. Each training window included two 224×224 standard-view images, language instructions, real-time 25-dimensional skeleton states, and future 32-step action states. Training occurred in two stages: an initial round of 6,000 parameter updates followed by a second round of 24,000 updates. Model selection was based exclusively on performance on the Franka validation set, with the weights from step 23,000 of the second phase chosen for evaluation.
This experiment compared four categories of mainstream robotic algorithms, covering current dominant technical routes. These included Diffusion Policy (a classic imitation learning algorithm), π0.5 and OpenVLA-OFT (vision-language-action models), Fast-WAM and LaWAM (world-action models), and Mirage and RoVi-Aug (cross-body visual adaptation schemes). All visual adaptation methods were built on Fast-WAM as the base policy. Every baseline algorithm was uniformly integrated with the target robot’s end-effector inverse kinematics solver, adhering to identical testing protocols and initial state standards. Only their native image processing and action output logic remained, and none could utilize skeleton images, centerline states, or body posture supervision information.
Results from the thousand-trial tests showed SkelWAM achieved an average task success rate of 43.3%, with a stable 95% Wilson confidence interval between 40.3% and 46.4%. Among all baselines, the best-performing model was RoVi-Aug equipped with end-effector inverse kinematics, achieving an average success rate of only 7.1%. SkelWAM led the baselines by 36.2 percentage points overall. In terms of specific robot performance, SkelWAM achieved the best results on nine of the ten robots. Success rates were 62.0% for UR10e and Rizon4, 59.0% for JAKA mini2 and Qing Yu Ke Ji A03, 56.0% for ARX-L5, 49.0% for Kinova Gen3, 32.0% for ViperX300, 22.0% for IIWA, and 19.0% for xArm7. Only the UR5e was slightly surpassed by a baseline; RoVi-Aug achieved 15.0%, marginally higher than SkelWAM’s 13.0%.
Model performance varies significantly depending on task type and robot category. In terms of tasks, the model excels at simple spatial interaction and object manipulation tasks, achieving success rates of 61.3% and 61.0%, respectively. Complex tasks, such as fixed-point targeting and long-horizon continuous execution, prove more difficult, with success rates dropping to 15.0% and 18.0%. Regarding robot types, industrial arms achieve an average success rate of 44.7%, lightweight arms 44.0%, and high-redundancy arms 38.0%. The Qing Yu Ke Ji A03 flexible arm achieves a single-machine success rate of 59.0%. Among high-redundancy seven-degree-of-freedom robotic arms, individual performance varies widely, with success rates ranging from 19.0% to 62.0%.
Ablation studies precisely validate the necessity of core modules. Replacing skeleton vision entirely with native color images from the source robot for training causes the model’s success rate to drop to zero. Removing the 15-dimensional body pose supervision while retaining only end-effector target constraints and the solver reduces the success rate to just 0.1%, demonstrating that skeleton vision and full-body pose supervision are the core pillars of the model’s capabilities. Removing the wrist camera view alone decreases overall success by only 0.3 percentage points, but performance fluctuates significantly across different tasks, with changes ranging from -6.7 and -14.7 to +22.5 and +8.0 across four task categories, indicating distinct adaptation characteristics for different scenarios. Removing the body constraint term from the solver while retaining complete model weights reduces overall success by 2.4 percentage points; object manipulation tasks see a drop of 8.7 percentage points, while long-horizon tasks improve by 3.0 percentage points, suggesting that body constraints adapt differently to various tasks.

Real-robot tests further verify the feasibility of cross-embodiment transfer, using completely different devices for training and deployment. The team collected 277 teleoperation data samples from a JAKA mini2 rigid robotic arm, splitting them into 249 training samples and 28 validation samples to construct a skeleton training window at 10 Hz. By fine-tuning only the last two layers of the action expert and video expert networks—completing 6,000 parameter updates—and evaluating solely based on the source device’s validation results, the model was ready for deployment. After training, all model weights were frozen, and the model was deployed directly onto the structurally distinct Qing Yu Ke Ji A03 cable-driven flexible arm to perform three desktop manipulation tasks: toy storage, ring stacking, and pin insertion. The target device required only updated calibration parameters and a curvature solver; no demonstration data from the A03 was needed.
Visual tracking in the real-robot tests relies on traditional computer vision algorithms. The team combines static region feature matching with the RANSAC algorithm, paired with forward-backward Lucas-Kanade optical flow calculation, supplemented by color thresholding to lock tool keypoints. This effectively addresses tracking issues such as occlusion and skeleton point drift. Through precise calibration of the task coordinate system, tool structure, and camera, numerical skeleton information is accurately aligned with dual-view camera feeds, standardizing the tool center point as the gripper contact center.

This part of the real-robot experiment serves only as a qualitative verification. Its core purpose is to demonstrate that policies trained on rigid arms can be transferred losslessly to cable-driven flexible arms with vastly different structures and execute tasks stably. The team did not collect quantitative success rates for the real-robot tests.
05 Conclusion
The core breakthrough of this work is the unification of robot visual observations, proprioceptive states, and action predictions into a single skeletal geometric representation. Observational information is generated via the skeleton, device states are defined by it, and action targets are output through it. This approach completely eliminates the rigid requirement for one-to-one joint structure correspondence in traditional cross-embodiment transfer. New robots no longer require repeated data collection or model retraining, enabling true zero-shot deployment.
For the field of embodied AI, this solution provides a universal, reusable technical interface. A model trained on a single Franka arm achieves state-of-the-art simulation performance across ten unknown robots. In real-world scenarios, policies trained on rigid arms can be directly adapted to flexible arms; cross-device deployment requires only swapping calibration parameters and kinematic solvers. The more diverse the robot embodiments and the richer the scenarios, the greater the savings in data collection and model iteration costs provided by this public skeletal representation, offering an efficient and viable new path for deploying general-purpose robotic intelligence strategies.
