Institute of Automation, Chinese Academy of Sciences and Amap Introduce DreamTrue, an Action-Faithful Robot World Model for Future Video Prediction

DreamTrue renders robot actions as visual conditions within the frame to supplement world models with failure demonstrations missing from demonstration data, and employs reinforcement learning based on a reward model trained via manual annotation. When an action fails, the model accurately predicts the corresponding outcome, reducing interaction defect rates by 48.12% to 6.25% on the AgiBot platform.

On October 8, the Institute of Automation, Chinese Academy of Sciences State Key Laboratory of Pattern Recognition and the Amap team published the paper "DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training." The paper has nine authors; Junyan Li and Ruizhi Li are co-first authors; Lue Fan, Assistant Professor at the State Key Laboratory of Pattern Recognition, and Zhaoxiang Zhang, Researcher at the Institute of Automation, Chinese Academy of Sciences, serve as corresponding authors; Yu Liu from Amap and Lue Fan jointly lead the project.

World models can serve as virtual simulation environments for robots, replacing some real-world trial-and-error testing. They must meet two core requirements: temporal alignment of actions and logical feedback from objects in response to those actions. Existing robot datasets contain paired actions and real-world video recordings, but training faces two major challenges.

First, most such videos lack precise annotations and complete camera parameters. Rendering actions into the footage causes misalignment with the original video, introducing bias into the foundational conditions the model learns from.

Second, datasets primarily feature successful demonstrations, almost entirely lacking failure scenarios such as grasping misses or object slippage.

This leads models to assume all actions succeed; even if a gripper passes over an object without touching it, the generated video still shows the object lifting. This bias inflates perceived success rates and misleads researchers who rely on world models to filter strategies.

DreamTrue addresses these pain points: it reverse-engineers camera calibration parameters from raw footage and uses counterfactual actions and a reward model to fill the gap of missing failure samples in training data.

01 Reverse-Engineering Camera Calibration from Footage

The method targets multi-view video prediction: it takes initial observations from multiple synchronized views, task instructions, and action sequences as input to predict subsequent multi-view videos.

The solution consists of three stages: offline geometric calibration based on existing recordings, supervised training in the first stage, and counterfactual post-training in the second. Among these steps, calibration is fundamental to ensuring precise alignment between action conditions and visual frames.

Offline geometric calibration specifically addresses the misalignment between rendered visuals and recorded footage. Leveraging recorded robot states and a set of initial camera parameters, the team renders the robot’s URDF model into the scene, projecting the robot’s geometry onto the pixel space of the video to establish a pixel-matching relationship between the rendered model and the actual robot in the footage. To build comprehensive constraints, in addition to this primary matching, the team introduces two supplementary constraints:

  • Matching the rendered robot with the recorded robot to verify the alignment between action conditions and visual frames;

  • Using static backgrounds across frames to constrain temporal consistency among cameras;

  • Using synchronized multi-view frames to constrain the relative geometric relationships between different cameras.

Match points are extracted by RoMaV2. Robot-related matches are restricted to the arm region, using masks generated by SAM 3 fine-tuned on a robot segmentation dataset to define the scope.

The three sets of match relationships participate in joint optimization. The solved parameters include camera intrinsics, extrinsics, distortion coefficients, and the mounting offset of the robotic arm. The match between the rendered and real robots serves to reduce reprojection error, while cross-frame background and cross-view scene points employ ray coplanarity constraints: the two camera rays observing the same point must remain coplanar with the line connecting the two camera centers. Joint optimization of these two types of constraints yields the final camera parameters and the robot’s mounting pose.

When processing real-robot datasets, the team grouped playback clips based on collection metadata, with grouping dimensions including the collector, camera, and robot ID. For each group, one playback was selected for calibration; the resulting parameters were then applied to evaluate other playbacks in the same group. Groups with poor alignment results received additional calibration. AgiBot and RoboMIND used fixed initialization; DROID had larger variations in camera extrinsics, using 32 groups of initialization schemes. RoboMIND and the self-developed Piper dataset additionally optimized for robotic arm installation offsets; RoboTwin and the migrated WidowX250 directly used the true geometric parameters from the simulator.

The team has open-sourced these calibration results, covering filtered real-robot playbacks from AgiBotWorld-Beta, DROID, and RoboMIND 2.0, totaling 153,666 segments and 1,660.31 hours of footage. Specifically, AgiBotWorld-Beta accounts for 1,225.60 hours, DROID for 239.73 hours, and RoboMIND 2.0 for 194.98 hours. Each record includes camera intrinsic parameters, extrinsic parameters, and distortion coefficients. Data from DROID and RoboMIND 2.0 additionally contain arm mounting offsets and calibration quality metrics.

02 Rendering Actions as Visual Conditions

After establishing a precise rendering geometric baseline through calibration, this method further addresses the challenge of unifying action adaptation across different robots and datasets. Since different robots have varying joint definitions, coordinate system rules, and end-effector representations, generic action encodings cannot be reused across scenarios. To solve this, the team innovatively transforms action information into visual representations. Leveraging the robot's URDF model and calibrated camera parameters, the complete action trajectory is rendered frame-by-frame and view-by-view into images. This aligns all action conditions to pixel space, effectively eliminating adaptation barriers for heterogeneous robots.

At every moment and observation angle, the model’s input conditions consist of five rendered information streams, comprehensively reconstructing the robot’s action state and spatial information:

First, RGB images of the robot, which concretize continuous actions into dynamic changes in the robot’s form within the image, intuitively presenting motion trajectories;

Second, depth images, which precisely output the spatial distance from various parts of the robot to the camera, preserving three-dimensional spatial scale information;

Third is the robot global mask, which clearly outlines all pixel regions occupied by the robot in the frame to achieve precise subject segmentation.

Fourth is the gripper open/close state, converted into background RGB intensity changes via linear mapping: the red channel represents the left hand and the green channel represents the right hand, quantifying the grasping action state.

Fifth is camera ray information, encoded as a dense Plücker image to precisely mark the spatial observation direction corresponding to each pixel in the frame.

These five condition streams are fed into the video diffusion Transformer model through the VACE branch. Rendered robot RGB images and normalized depth images undergo feature encoding via a pre-trained video VAE; mask images and Plücker ray images are input into a 3D convolutional geometric encoder. After multi-stream features are fused and concatenated to form the context input for the VACE branch, they are injected into the corresponding layers of the DiT network in residual form. For multi-view tasks, the model concatenates latent variables from different views along the width dimension in a fixed order, synchronously matches action features to the corresponding layout, and uses initial multi-view observation frames as reference frames to ensure temporal and spatial consistency.

The model generator is initialized based on Wan2.1-VACE-14B, supporting processing of 101 frames and three synchronized observation views in a single pass, with a single-view resolution of 240×320. The model takes the first-frame RGB images and task instructions from each view as base conditions to predict continuous video frames for the subsequent 100 frames. Regarding camera configuration, the DROID dataset uses a combination of two external cameras plus one gripper-mounted camera, while other datasets use one main camera plus two wrist-mounted cameras.

The first training phase focuses on the DiT network, the LoRA branch of VACE, and the 3D convolutional geometric encoder, with approximately 700 million trainable parameters and a LoRA rank set to 128. This phase uses paired real-world future videos as supervision signals. By inputting initial observation frames, task instructions, and standardized action conditions, it employs flow matching algorithms to gradually denoise real multi-view video latents. The loss function is the squared error between predicted velocity and true velocity, guiding the model to learn the correspondence between actions and visual changes to accurately generate frames following the robot's motion logic.

Training data for this phase covers four major datasets: AgiBotWorld-Beta, DROID, RoboMIND 2.0, and RoboTwin 2.0. After strict filtering, the cumulative dataset includes 2,232 hours of multi-view trajectory data covering five types of robotic arms, compatible with both real-robot collection and simulation scenarios, ensuring the model's generalization and robustness.

03 Using Self-Constructed Failure Samples for Post-Training

The first stage of training relies entirely on successful interaction samples from real-world datasets. Consequently, the model learns only the visual logic associated with effective actions and lacks experience with failure scenarios, leading to a bias where it predicts success even when actions are erroneous. To address this, the research team introduced a second stage of counterfactual post-training. Without collecting additional data from physical robots, they autonomously constructed a large volume of failure action samples and trained a dedicated reward model to accurately score and correct errors in the videos generated by the model.

The team built these counterfactual failure actions based on successfully demonstrated trajectories recorded in reality. Specifically, they applied SE(3) spatial perturbations to the end-effector pose of the demonstration trajectory, interpolating from a fixed initial pose to an abnormal endpoint after perturbation, and then solved for the corresponding robotic arm joint sequences using inverse kinematics. Throughout this process, the initial observation images and task instructions remained completely unchanged; only the action sequence was replaced. All invalid samples that did not meet kinematic feasibility were filtered out. Leveraging a fixed initial scene, a single original video could generate multiple branches with different actions and outcomes, enabling the model to learn how different actions lead to different endings within the same scene, thereby filling the data gap for failure scenarios.

Because the autonomously constructed failure actions lacked corresponding real-world future videos as ground truth, conventional loss functions could not be used to determine correctness. The team shifted their optimization strategy: instead of comparing against visual ground truth, they detected observable physical and interaction defects in the generated videos to specifically correct model biases. Researchers generated videos based on several mainstream world models under conditions of both standard recorded actions and constructed counterfactual actions, and manually annotated three categories of typical defects uniformly.

These three defect dimensions cover the core issues in model generation:

First, robot body defects, including ghosting, disappearance, structural disarray, and material anomalies of the robot itself;

Second, object defects, including the disappearance, deformation, appearance confusion, and texture or material distortion of target objects;

Third, interaction defects, including misaligned contact relationships, inconsistent motion logic, and interaction behaviors that violate physical laws.

This round of manual annotation covered a total of 44,900 raw video clips and 30,400 defect labels. After filtering and cleaning, 36,493 valid data samples were retained for reward model training and validation. The training set contains 31,600 samples and the validation set contains 4,893 samples, split strictly according to the original playback IDs. To ensure the generalization capability of the reward model, the annotated corpus includes not only videos generated by this method but also content produced by seven external world models: DreamDojo, Ctrl-World, EnerVerse-AC, Genie Envisioner, GE-Sim 2.0, Cosmos Predict 2.5, and IRASim. Additionally, the Bridge dataset was introduced to assist with training and validation.

The team fine-tuned the Qwen3.5-9B vision-language model using high-quality annotated data, enabling it to simultaneously output defect probabilities across three dimensions: body, object, and interaction. The study used the negative value of defect probability as a reward score; lower visual defect probabilities resulted in higher reward scores for the model, thereby guiding it to generate video frames that align with physical laws and accurate interaction logic.

The second stage of post-training iteratively optimizes the Stage 1 pre-trained model. The training process inputs both real recorded actions and self-constructed counterfactual actions; the generator samples future videos under these two action conditions, and a frozen reward model scores them. For videos generated from real recorded actions, an additional PSNR reward constraint is applied by comparing them against ground truth frames.

All dimensional rewards undergo intra-group normalization via GDPO, followed by batch-level normalization after weighted summation to compute the final advantage value, with generator parameters updated using the DiffusionNFT algorithm. This stage employs a freezing strategy: the base generator, geometric encoder, and reward model parameters are fixed, while only the DiT network and VACE's LoRA branches are updated. The weights for all three defect rewards are set to 1, the PSNR reward weight corresponding to recorded actions is set to 0.5, and the training learning rate is set to 5×10⁻⁶.

04 Predicted Performance and Strategy Evaluation

The experiments evaluated action following, visual quality, and interaction rationality on AgiBot; cross-body and cross-environment performance on DROID, RoboMIND 2.0, and RoboTwin 2.0; and separately measured policy result evaluation, geometric calibration, and reward model effectiveness.

The evaluation includes the full proposed model, a one-stage model trained without reinforcement learning, and four baselines—DreamDojo, GE-Sim 2.0, Genie Envisioner, and EnerVerse-AC—each using its native conditional interface. Action following is measured using nDTW from EWMBench; image quality is assessed via PSNR, SSIM, and LPIPS, along with WorldArena’s EWMScore-P, which aggregates 15 metrics across six quality dimensions. Interaction plausibility is evaluated by three annotators independently, with defect rates calculated via majority voting. AgiBot’s counterfactual evaluation uses 160 conditions covering 54 tasks.

With recorded actions, the full model achieves an nDTW of 0.8772, outperforming EnerVerse-AC (0.8058), Genie Envisioner (0.8124), GE-Sim 2.0 (0.8063), and DreamDojo (0.6578). Under counterfactual actions, it reaches 0.8831. In terms of visual quality, the model scores a PSNR of 23.06, SSIM of 0.936, and LPIPS of 0.094. Training with counterfactual data reduces interaction defect rates from 48.12% to 6.25% and object defect rates from 31.88% to 3.12%, while action-following nDTW improves slightly from 0.8711 to 0.8772. For comparison, GE-Sim 2.0 reports interaction and object defect rates of 7.50% and 3.75%, respectively, on the same metrics.

The reduction in defect rates corresponds to clearer visual outcomes: the model no longer renders missed grasps as successful ones. When the gripper misses, it keeps objects on the table, whereas comparative models often make objects disappear, duplicate, or float upward. By fixing the initial scene and varying only the actions, the model generates distinct consequences based on the input. Objects caught by the gripper move with it, while uncaught items remain stationary; even in single-arm operations where only one hand succeeds, the model correctly distinguishes between them.

In the World Model track of the AgiBot World Challenge 2026, this method achieved a comprehensive score of 0.829, an action following score of 0.9651, and a visual quality score of 0.6246, ranking first among participating teams. In cross-embodiment evaluation, the same set of weights yielded a PSNR of 22.95 on DROID, which is higher than Ctrl-World's 22.00, and an LPIPS of 0.073, which is lower than its 0.162; RoboMIND 2.0 achieved a PSNR of 23.84, and RoboTwin 2.0 achieved a PSNR of 27.35. When the WidowX250, which was not involved in training, was placed into RoboTwin 2.0 and then transferred to Piper data collected from unseen real-world scenes, it could still provide predictions without fine-tuning.

Policy evaluation used three two-handed tasks from RoboTwin 2.0—beat_block_hammer, handover_block, and place_dual_shoes—on the Aloha-AgileX embodiment. The team tested four open-source VLA policies (EventVLA, π0.5, X-VLA, and starVLA) across 60 fixed scenarios, generating 240 action replays and 24 task-policy-setup combinations. The world model then predicted outcomes from these actions, which annotators judged for success or failure. The full model’s predicted success rate deviated from real-robot simulation averages by 7.08 percentage points, compared to 12.92 percentage points for Ctrl-World. Overall bias was +0.42 percentage points, with a rank correlation of 0.937. In the low-success-rate range, Ctrl-World’s false positives dropped from 26 to 12, improving precision from 76.15% to 87.63% and accuracy to 90.42%.

Geometric calibration performance was measured using area-weighted IoU between rendered robots and SAM 3 masks. Compared to dataset-provided calibrations, this method improved average IoU by 0.228 across 130,182 AgiBot replays and by 0.212 across 63,061 DROID replays, showing improvement in 97.1% and 91.6% of cases, respectively. On DROID, using only RGB inputs, it outperformed PointWorld—which relies on stereo depth—in 75.5% of 38,356 replays. Integrating this calibration into prediction boosted PSNR during inference alone; when also applied during training, PSNR increased by 1.28 dB for AgiBot and 1.30 dB for DROID, with synchronous reduction in end-effector position synchronization errors. On 4,893 held-out video segments, the reward model achieved an average discrimination accuracy of 86.33% for three defect types and a macro-average F1 of 74.21%, with embodiment-specific accuracy at 96.69%, and object and interaction accuracies at 79.24% and 83.08%, respectively.