Five Universities Release AffordanceWAM to Bridge Human and Robot Training Data
The core breakthrough of AffordanceWAM is incorporating affordances—the object’s interactive locations and methods—into the explicit prediction targets of world-action models. This resolves a long-standing industry pain point: adapting first-person human videos for robot training. It transforms human data, which previously risked negative transfer, into scalable, effective supervision signals.
On September 16, a joint team from The University of Hong Kong (Xiang Gang Da Xue), Hong Kong University of Science and Technology (Guangzhou) (Xiang Gang Ke Ji Da Xue ( Guang Zhou )), Peking University (Bei Jing Da Xue), The Chinese University of Hong Kong (Xiang Gang Zhong Wen Da Xue), National University of Singapore (Xin Jia Po Guo Li Da Xue), and Knowin AI released their research paper, AffordanceWAM: Affordance-Aware Joint World–Action Modeling for Robot Manipulation. The paper features 14 authors, with Jiadi You and Qize Yu as co-first authors. Associate Professor Qi Xiao Juan from the Department of Electrical and Electronic Engineering at The University of Hong Kong (Xiang Gang Da Xue) and Assistant Professor Chen Ying Cong from the Artificial Intelligence Domain at Hong Kong University of Science and Technology (Guangzhou) (Xiang Gang Ke Ji Da Xue ( Guang Zhou )) serve as corresponding authors.
World-Action Models (WAMs) introduce temporal video prediction capabilities into action generation logic, outputting control commands based on an understanding of future scene states. Compared to concrete joint actions that are difficult to transfer across different embodiments, future states such as visual changes and interaction trends are more universal, naturally facilitating transfer across diverse agents. However, when training with human operation videos, the industry has long faced two major misalignment issues: human videos lack robot joint action annotations, and the differences in human-machine perspectives, physical structures, and operational rhythms are significant. Directly mixing these data sources easily interferes with model training.
To address this, the team innovatively defines affordances—the physical attributes of an object’s interactive regions and operable methods—as one of the future world states the model must predict, achieving unified supervision for human and machine video data. Based on this design, human videos can directly participate in world-model training without requiring action annotations or pre-computed human-to-robot action retargeting, providing effective supervision for robots learning scene interaction logic.
01 Semantic Misalignment Between Human Videos and Actions
First-person human videos fully present scene information, including hand movement trajectories, object placement positions, and interaction contact points, but completely lack the joint control instructions required by robots. Meanwhile, there are fundamental differences between human hands and robotic arms in degrees of freedom, motion structure, and operational rhythm. For the same interaction task, the action coordinates and temporal logic differ entirely between humans and machines. If human videos are directly mixed into robot training datasets, the model cannot obtain effective action supervision; instead, it suffers interference due to misaligned visual representations and motion logic, degrading training performance.
Previous industry solutions could only patch shortcomings unidirectionally. One approach relied on human-to-robot action retargeting, forcibly converting human trajectories into robot action sequences. Training effectiveness depended entirely on the accuracy of action reconstruction and cross-embodiment alignment, leaving very low tolerance for error. Another approach focused on visual adaptation, aligning human and machine observational views through perspective and appearance synthesis. However, this only resolved visual discrepancies, failing to fix the underlying misalignment in interaction logic. Unaligned scene information remained as training noise. Both approaches required high matching between human and machine data, resulting in high adaptation costs and poor generalization.
World-Action Models offer a new technical path for this challenge. By sharing backbone networks for video prediction and action generation, WAMs rely on common future scene states to infer actions. Scene temporal features are easier to transfer across embodiments than concrete joint actions. However, most existing WAMs only predict future frames or visual latent variables, modeling only changes in appearance, perspective, and background motion. Key information such as object interaction regions and operable points remains hidden in implicit features, unable to form effective, stable supervision signals.
The core innovation of this study is transforming affordances from implicit features into explicit prediction targets. Affordances represent the objective, interactive attributes of objects, unaffected by differences between human and robotic agents or by viewing angles and pacing. They provide a unified interaction logic shared by both humans and machines. By integrating them into the predictive framework of world models, human videos and robot videos share a common supervisory dimension, enabling efficient training for embodied AI without cumbersome data preprocessing or motion adaptation.
02 Embedding Affordances into Future World Models
AffordanceWAM decomposes affordances into two parallel branches: scalar affordance score maps and affordance heatmaps. These dual representations complement each other to capture scene interaction information, establishing a unified training pathway that connects human video data with robot data.
A scalar affordance map is a pixel-level score map with the same dimensions as the original image. Pixel values are normalized to the 0–1 range, allowing frame-by-frame localization of task-relevant objects and interactive regions. This representation defines only the 'interaction location,' without constraining specific grasping postures or execution methods. It naturally avoids discrepancies in body structure and movement habits between human hands and robotic arms, allowing human videos and robot videos to share a unified interaction semantics.
The scalar affordance is modeled through a dedicated, trainable encoder-decoder architecture. The encoder compresses multi-frame future affordance score maps into latent variables aligned with temporal groupings of the footage. The decoder then reconstructs the score maps, trained via frame-by-frame mean squared error supervision to precisely preserve spatial interaction position information. However, pure score maps lack visual attribution; isolated interactive regions cannot be linked to corresponding objects, potentially causing model recognition confusion. To address this, the team added an affordance heatmap branch. Using a fixed rendering method, the abstract score map is overlaid onto the original footage. A frozen video encoder then encodes this into latent variables, strictly aligning temporally with the visual features. This deeply binds interactive regions to their parent objects, supplementing the scene's appearance and semantic attribution information.

The model applies flow-matching constraints uniformly across three future states: the visual frame, the scalar affordance, and the affordance heatmap. Independent noise is injected into each branch, which are then fused using weighted coefficients to combine temporal flow loss with encoder-decoder reconstruction loss, forming a complete world model supervisory system. The training strategy distinguishes precisely between data sources: human videos participate only in the three-way world state prediction supervision, focusing on learning general scene interaction patterns. Robot trajectory data, in addition to the world loss, includes extra action supervision. Relying on this dual-track mechanism, affordances become the core universal representation bridging human data and real-robot data, completely resolving the misalignment issues in training these two types of data.
03 Dual-Expert Architecture and Unidirectional Attention
AffordanceWAM employs a dual-branch diffusion Transformer expert architecture, with two groups of experts, each consisting of 30 layers, operating in decoupled collaboration. The world expert is initialized using the pre-trained video diffusion model Wan2.2-TI2V-5B, features a hidden dimension of 3,072, and is responsible for jointly predicting three world states: future multi-frame images, scalar affordances, and affordance heatmaps. The action expert, with a hidden dimension of 1,024 and a feedforward dimension of 4,096, performs inference based on the hidden features from corresponding layers of the world expert, ultimately outputting continuous robot control actions across 12 steps, each with 7 dimensions.
The two experts achieve feature interaction through a masked joint self-attention mechanism, while strict unidirectional flow constraints prevent interference between branches. Image and scalar affordance features can exchange information bidirectionally; however, action features can only passively read image and affordance information without influencing world modeling. Heatmap features also maintain independent, controlled exchange. The core masking mechanism blocks two specific interference paths: first, it prevents action features from flowing back to affect world prediction, ensuring pure scene understanding; second, it completely blocks heatmap features from entering the action branch, eliminating reasoning shortcuts driven by visual appearance preferences.
This design ensures that heatmaps contribute to shaping general world features solely through gradients. The action expert cannot rely on image textures as a shortcut but must instead focus on learning more stable, cross-embodiment scalar affordance interaction logic, structurally enhancing the model’s generalization capability in real-world scenarios.

Model training consists of two phases. Phase one is joint pre-training on heterogeneous human-robot data: human videos update only the world expert parameters, supervising the three future state predictions; robot trajectory data adds action supervision on top of this. The process includes warm-up training for the pure world model, with action supervision introduced progressively to ensure stable convergence. Phase two involves specialized fine-tuning: the complete world model weights are frozen, and only the action expert is optimized using robot vision-language-action data over 10,000 iterations with a learning rate of 10⁻⁵. A mixed sampling strategy is used for batch training: 75% of samples use noisy world features, while 25% use fully predicted world features, balancing generation robustness with final execution accuracy.
Training data is sourced entirely from public high-quality datasets covering massive natural human-robot interaction and robot operation scenarios. Human video materials are drawn from Ego4D-FHO, EPIC-KITCHENS VISOR, HOI4D, H2O, and MECCANO; robot trajectory data includes DROID and RH20T from RoboInter, as well as BridgeData V2, InternData-A1, and RoboCasa365. The study requires no additional training of dedicated annotation models; instead, it directly reuses native object and interaction region annotations from the datasets. Affordance score maps are generated via Gaussian kernel rasterization, with per-frame peak normalization. Missing annotation regions are filled with at most two frames of interpolation, and occluded scenes retain only valid features from visible areas, ensuring precise annotations that align with real interaction states.
Inference adopts a two-stage generation-and-correction paradigm. Four sets of latent variables—images, scalar affordances, affordance heatmaps, and actions—are initialized with noise. The dual experts interact layer-by-layer over 35 iterations to generate world states and action sequences. After the initial generation, the world model output is fixed, and the action sequence undergoes refined optimization with 8 correction steps and an intensity of 0.25, producing a stable, executable action sequence. The entire training was conducted on 64 H200 GPUs using the AdamW optimizer with BF16 precision. Each training segment totals 17 frames, comprising 5 frames of observation input and 12 frames of future prediction, at a frame rate of 10 fps and an image resolution of 256×320.
04 Results on Simulation and Real Robots
The team validated this approach on two simulation benchmarks and one real-robot platform. On RoboCasa, AffordanceWAM achieved an average success rate of 69.4%, surpassing the previous best Cosmos Policy at 67.1% by 2.3 percentage points; π0 and GR00T-N1.5 scored 62.5% and 64.1%, respectively. In the CALVIN ABC→D test, the success rates for completing one to five consecutive tasks were 96.1%, 92.8%, 85.5%, 77.8%, and 69.5%, with an average completed sequence length of 4.22, outperforming the previous strongest baseline VPP's 4.01.
A set of key controlled experiments fixed the volume of robot-side supervision data, adjusting only whether affordance prediction was enabled and whether real-human video was included. When trained solely on robot data without affordance prediction, RoboCasa scored 57.3% and CALVIN reached 3.75; adding real-human video to the same dataset caused both metrics to drop to 55.1% and 3.68, indicating that real-human video then hindered model performance. After enabling affordance prediction, training with only robot data yielded scores of 63.7% and 3.91; introducing the same batch of real-human video further improved both metrics to 69.4% and 4.22. On RoboCasa, the performance gap attributable to the presence or absence of real-human video reached 7.9 percentage points.
With the robot data and training pipeline held constant, gradually adding real-person videos with affordance annotations caused RoboCasa's success rate to increase monotonically from 63.7% to 69.4%. The first 60% of real-person videos contributed only a 1.5 percentage point improvement, while the final 40% delivered a 4.2 percentage point gain, accounting for approximately 74% of the total improvement.
Ablation experiments quantify the contribution of each module. Removing the affordance heatmap branch caused RoboCasa scores to drop from 69.4% to 65.4%, and CALVIN from 4.22 to 3.99. Releasing the attention from the heatmap to the action branch, allowing the action expert to read heatmap features, further reduced scores to 63.5% and 3.85. Changing the expert coupling method also resulted in performance degradation: performing feature fusion only at the end of the network yielded 65.9% and 4.00; generating a complete world prediction first and then running the action expert separately resulted in 64.5% and 3.90; canceling exposure to predicted world features during second-stage training resulted in 66.6% and 4.06. All control experiments used three random seeds, with 1,200 evaluation episodes run per seed for RoboCasa.

Real-robot experiments compared AffordanceWAM and Cosmos Policy under identical conditions. Both models were fine-tuned using 50 demonstration samples per task, with each task evaluated over 15 trials. The five basic tasks required picking up target objects based on category, color, and shape, then placing them. AffordanceWAM achieved an average success rate of 74.7%, while Cosmos Policy reached 42.7%. In three more complex tasks, the success rate for closing drawers improved from 46.7% to 80.0%; fruit-picking results were graded at 86.7%, 73.3%, and 53.3%; and flower-arranging performance was 20.0% versus 6.7%.
05 Final Thoughts
The most significant change in this study is establishing a shared prediction target for human and robot videos. Previously, leveraging human operation data required either retargeting human motions to robot joint actions or aligning the visual scenes from both domains. AffordanceWAM unifies these two data types by binding them to affordance prediction. This allows human videos to provide supervisory signals for world models without requiring motion annotations or prior motion retargeting.
It leaves interfaces for further expansion. Affordance annotations can be generated directly from existing object and interaction annotations. The larger the volume of human video data, the more comprehensive the coverage of affordance annotations, and the greater the performance gain for the robot model. In the paper's experiments, the final 40% of human videos contributed approximately 74% of the performance improvement. To extend this approach to more task scenarios, data that allows for the annotation of interaction regions serves as the entry point for the entire pipeline.
