Shanghai Jiao Tong University and LandSpace Team Release DynaForge to Turn Failed Demonstrations into Training Data

DynaForge eliminates the need for manual, frame-by-frame collection of dynamic manipulation demonstrations. Robots can autonomously trial runs in simulation environments, learning correction values from failed episodes to transform originally failed trajectories into effective demonstrations for policy training.

On September 22, the School of Automation and Perception at Shang Hai Jiao Tong Da Xue, in collaboration with a team from Lan Jian Hang Tian Kong Jian Ke Ji Gu Fen You Xian Gong Si, published the paper "DynaForge: Planning-Guided Residual Learning for Dynamic Manipulation Demonstration Generation." The paper has four authors, with Yiyang Jin and Yu Zheng as co-first authors, and Professor Wang He Sheng of the School of Automation and Perception at Shang Hai Jiao Tong Da Xue serving as the corresponding author.

Dynamic manipulation is an unavoidable task type for robot deployment. Objects may roll, slide, or rotate, requiring robots to coordinate perception, motion, and contact within extremely short time windows. Existing large-scale datasets and automated demonstration pipelines mostly target static or quasi-static scenarios; demonstrations reflecting robot-object interaction remain scarce, making it difficult for imitation learning and vision-language-action models to acquire sufficient dynamic interaction data. Manual teleoperation struggles to complete operations stably within brief contact windows, resulting in high collection costs. Motion retargeting from human videos faces issues such as morphological differences and misaligned contact timing. While recent automated solutions use state machines, goal updates, and trajectory reconstruction to adapt to dynamic scenes—providing geometric guidance through planning and retargeting—the execution end cannot autonomously learn corrections from multiple failures, showing significant shortcomings at critical contact boundaries. By treating motion planning as a structured prior, where planning handles global motion construction and the learning module assumes execution-level correction, this decomposition approach led to the proposal of DynaForge.

01 Contact Boundaries Determine Success

The contact boundary is precisely the stage where dynamic operations are most prone to failure. Motion planning can solve for a collision-free global path but struggles to determine whether the robotic arm should close its gripper early or late when an object rolls to a specific position. As objects move continuously, the window for effective contact is extremely narrow; even minor deviations in planning during this period can cause the entire task trajectory to fail.

Another class of approaches forces robot actions to align with preset object motion trajectories, simplifying contact dynamics. Although tasks can be completed, robots do not dynamically adjust actions based on real-time object movement, limiting the reference value of generated experience samples for policy training.

The team decomposed the problem into two layers: global collision-free motion is handled by motion planning, while phases near contact that require real-time response rely on learned residuals for correction. The entire task is segmented by phase, with each phase corresponding to a target in the object's coordinate system, success criteria, control mode, and residual switch.

The study set up nine simulation tasks covering dynamic grasping, catching, tapping, placing, and inserting, involving rolling objects, moving containers, and rotating fixtures, with object speeds ranging from near-static to 1.5 meters per second. The simulation environment used Isaac Lab, with the Franka Research 3 robotic arm, and scene assets sourced from RoboTwin and DOM.

During evaluation, each task uses three sets of random seeds, with 100 episodes per set. Demonstration generation success rate refers to the proportion of episodes that successfully produce valid demonstrations out of the total attempts; all subsequent generation metrics follow this definition.

02 Planning as Prior, Residuals for Contact

The action executed by DynaForge at each control step is composed of two parts: a nominal action provided by the planner and a correction added by a residual policy. These are summed and passed through an action limiter.

The planner operates in two modes based on phase. During the stable free-space phase, the planner solves for the target only once at the beginning of the phase, generating a collision-free trajectory that is executed continuously until the success criteria for that phase are met. Global replanning is triggered only when switching phases. If the system enters an object-motion or contact-sensitive phase, the target updates in real time with the object, and inverse kinematics is solved at every step based on the object's latest pose.

The target pose remains fixed in the object coordinate system. The end-effector target pose in the world coordinate system is derived from the object's real-time pose. Once the object moves, the end-effector target migrates synchronously, and the robotic arm continues to follow the object's new position.

Inverse kinematics can only solve for geometric targets and cannot handle temporal and interaction errors present in contact regions. This part is taken over by the residual policy during the corresponding phase. The policy receives a short history of robot-object states and planned actions, samples from a Gaussian distribution, and outputs residual actions. Residuals are active only in phases requiring dynamic response; in other phases, they are disabled via a gating mechanism. Exploration noise is also removed during evaluation and formal demonstration generation.

The training process is conducted entirely within a simulation environment. The residual policy accesses the full state information of both the robot and the object. This type of privileged information is available only during the training phase and cannot be directly obtained in real-world deployment scenarios.

03 Learning Only from Half-Successful, Half-Failing Groups

The residual policy is trained using reinforcement learning. The model does not employ step-by-step dense rewards; instead, it determines task success or failure only at the end of each episode. It supplements this with partial rewards based on task progress checkpoints. The return for a single episode is determined by whether the task was successful and which checkpoint was reached in the event of a failure.

If environmental conditions are sampled randomly, task difficulty and the effectiveness of residual corrections become confounded, making it impossible to attribute differences within groups. To address this, the research team categorized difficulty levels based on parameters such as object velocity and disturbance magnitude, assigning an equal number of environments to each level. Environments within the same group share initialization and disturbance parameters but have independent action sampling, ensuring that differences within the group stem solely from the execution process itself.

Gradient updates are performed only on specific sample groups. If all samples in a group succeed, it indicates that the current conditions have been mastered, leaving no failed cases requiring correction. Conversely, if all samples fail, the difficulty exceeds the model's capability range, preventing a meaningful comparison between success and failure. Both types of homogeneous groups are skipped. Only groups containing both successful and failed samples participate in gradient calculations. If a batch contains no qualifying groups, the update for that round is skipped.

As the policy's capabilities improve, these mixed-success/failure groups automatically migrate toward higher-difficulty conditions, keeping the learning focused on the model's current performance boundaries. The study does not use manually preset difficulty sequences; instead, difficulty increases naturally through the filtering mechanism. This approach is referred to as implicit curriculum. The optimization algorithm employs DAPO asymmetric truncation, with a lower bound of 0.20 and an upper bound of 0.28.

The study also documented the entire progression of the capability frontier. Each task generated a hierarchical success-rate heatmap, accompanied by two training curves corresponding to the implicit curriculum and standard GRPO, respectively. The heatmaps visually demonstrate that as training progresses, success rates spread toward higher difficulty levels, continuously expanding the capability frontier.

04 Generation Results and Real-Robot Transfer

First, let's look at the demonstration generation results. Across nine simulation tasks, planning priors alone yielded an average success rate of 41.30%. With the addition of a residual strategy, this improved to 78.37%, representing an average gain of 37.07 percentage points. The improvements varied significantly across tasks: lemon grasping increased by 12.00 percentage points, while bottle grasping saw the largest jump of 57.67 percentage points; success rates for tasks such as toy car manipulation, alarm clock interaction, and block stacking exceeded ninety percent. For the most challenging task—pin insertion—the success rate rose from 7.67% to 28.00%.

Two baseline approaches performed worse. DynamicVLA’s state-machine collector achieved an average success rate of 29.07% across its five supported tasks; the remaining four tasks fell outside the state machine’s coverage and were marked as not applicable (N/A). DynaMimicGen averaged only 13.41% across all nine tasks.

The team estimated preparation costs for five common tasks shared by each method. DynaForge required approximately 50 minutes of human effort and 6.5 hours of machine time, comprising 2.5 hours for residual training and 4 hours for data collection. By comparison, DynamicVLA required an estimated 100 minutes of human effort and 26.9 hours of machine time; DynaMimicGen required 52.5 minutes of human effort and 25.4 hours of machine time; and the baseline planning-only approach required 75 minutes of human effort and 10.4 hours of machine time. The team explicitly noted that these human-effort estimates are hypothetical valuations rather than measured results.

The value of the generated data was validated through downstream policy training. Each task used 800 demonstration samples to train a 3D Diffusion Policy. Two experiments maintained identical representation methods, preprocessing pipelines, and training steps, differing only in the data source. Policies trained on the DynaForge dataset achieved an average success rate of 49.11% across the nine tasks, whereas those trained on the DOMINO dataset averaged just 7.07%. The gap was most pronounced in the alarm clock and toy car tasks, while pin insertion remained high-difficulty for both datasets.

An ablation study isolated the impact of implicit curriculum learning. Across four tasks, using planning alone yielded an average success rate of 37.75%; standard GRPO training reached 63.25%; and enabling the implicit curriculum raised it to 82.75%. Compared to standard GRPO, the implicit curriculum improved performance by 1, 5, 24, and 48 percentage points across the four tasks, respectively, while requiring only 0.73 times the optimization steps and 0.62 times the training time.

The research further transferred to real-robot environments. The team directly deployed the 3D Diffusion Policy trained in simulation for three physical experiments: picking up and lifting a soda can rolling in from the right, pressing down on a continuously rotating alarm clock, and placing a banana into a bowl on a turntable. Each task and data source combination was tested ten times. Policies trained on DynaForge data succeeded 3, 6, and 5 times, corresponding to success rates of 30%, 60%, and 50%. Policies trained on DOMINO data succeeded 0, 0, and 1 times, corresponding to success rates of 0%, 0%, and 10%.

Here, the division of labor must be clarified. DynaForge is a front-end data generation tool that produces dynamic interaction demonstrations in bulk within simulation environments. The diffusion policy trained on these demonstrations runs on the physical robot; the DynaForge generator itself is not deployed on the robot hardware.

05 Conclusion

DynaForge uses motion planning as a prior and relies on learning modules to correct timing errors at contact boundaries, transforming originally failed trajectories into valid demonstration samples. The core value of this approach lies in enabling the batch automatic generation of dynamic operation demonstrations, eliminating the need for manual collection one by one during extremely short contact windows.

For embodied AI researchers, this constitutes a reusable data pipeline. The dynamic operation demonstration generation rate increased from 41.30% to 78.37%; policies trained with generated data achieved an average success rate improvement from 7.07% to 49.11% in simulation, and attained success rates between 30% and 60% in real-robot tests. The more dynamic skills that need to be covered, the more significant the acquisition cost savings provided by this automated generation pipeline.