StageCraft Authors on Why VLA Generalization Fails in Changing Environments

The IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) was held in Pittsburgh, USA, from September 27 to October 1.

This year’s IROS accepted nearly 1,900 papers, organized into more than 300 session tracks.

As one of the two flagship top-tier conferences in robotics alongside ICRA, IROS covers a comprehensive range of research directions. It encompasses core fields such as robot structural design, motion control, environmental perception, robotic AI, dexterous manipulation, autonomous driving, and human-robot interaction, gathering cutting-edge global research outcomes.

Held entirely in Pittsburgh, the multi-day conference featured a tight schedule with back-to-back presentations across various venues. Even with full-day immersive attendance, we could only sample a fraction of the frontier work presented.

One notable session focused on StageCraft.

This work, completed by the team from Ya Li Sang Na Zhou Li Da Xue, is titled "StageCraft: Execution Aware Mitigation of Distractor and Obstruction Failures in VLA Models." The Chinese meaning is roughly "Execution-stage mitigation of interference and obstruction failures in VLA models." The paper was officially released in March and has been accepted by the IEEE/RSJ International Conference on Intelligent Robots and Systems 2026.

This research focuses on the generalization boundaries of Vision-Language-Action (VLA) models. While mainstream VLA models, trained on large-scale datasets, demonstrate strong cross-task and cross-scenario capabilities, their task execution success rates drop significantly in real-world robotic workbench scenarios when encountering out-of-distribution distractors or physical occlusions.

Such performance failures are common in real-world robotics deployment.

In most cases, the VLA model itself is capable of completing the target task; the failure stems from environmental interference rather than insufficient model capability. Traditional optimization approaches have clear limitations: fixing generalization issues by re-collecting data and fine-tuning the model is costly and inefficient. Moreover, many real-world scenarios cannot replicate the original training data, making these optimization efforts difficult to implement.

Unlike mainstream methods that stack optimization strategies within the training pipeline, StageCraft takes a different approach. It requires no retraining of the model and no modification of policy parameters; instead, it resolves failures by adjusting the initial state of the environment. The system first collects task execution data from the robot across various distractor scenarios, uses a Vision-Language Model (VLM) for reasoning and analysis to precisely identify key objects causing task failure, and finally removes only the minimal set of distractors necessary to eliminate environmental interference.

Real-robot experiments show that in complex scenes with distractors, this method increases the average task success rate by approximately 40 percentage points. Tests covered three typical robotic manipulation tasks: stacking cups, arranging plates, and placing blocks into bowls, demonstrating significant improvements.

This research was independently conducted by the team at Ya Li Sang Na Zhou Li Da Xue, with Kartikay Milind Pangaonkar, Prabin Kumar Rath, and Omkar Patil as co-first authors, under the supervision of Nakul Gopalan, Assistant Professor in the Department of Computer Science at Ya Li Sang Na Zhou Li Da Xue.

Following the presentation, we held an in-depth conversation with Prabin Rath, one of the paper’s co-first authors.

Below is the transcript of the dialogue between 42HOW Robotics and Prabin Rath, lightly edited for clarity.

From Embodiment Transfer to Environmental Variation

42HOW Robotics: Could you briefly introduce your academic background?

Prabin Rath: Hello everyone, I’m Prabin Rath. I am currently a second-year Ph.D. student at Ya Li Sang Na Zhou Li Da Xue, conducting research under the guidance of Professor Nakul Gopalan.

My work primarily focuses on holistically enhancing robots’ learning capabilities. Vision-Language-Action (VLA) models are a key class of robotic learning models used in our research.

Previously, my research focused on transferring policies across different robot embodiments—enabling these models to generalize from one robot to another without retraining. Recently, I have also been investigating policy memory: how robot models can operate effectively over longer contexts, retain past events, and make informed decisions based on that history. Current VLA models do not yet possess these capabilities, and my research explores how to integrate them into state-of-the-art models.

42HOW Robotics: Compared to your previous research on robot behavior and embodiment, what is the biggest challenge in shifting to environmental change studies like StageCraft?

Prabin Rath: Robot embodiment transfer problems are quite different from the issues addressed in this work.

Embodiment transfer is harder in some respects. There are not many robot embodiments available in industry, and the types of robots purchasable on the market are limited. If we train a policy on only one robot, it may fail to generalize to other robots.

The environmental issues handled by StageCraft are almost independent directions from embodiment issues, but they present equally significant challenges. We lack sufficient robot data for training, and when a policy is moved from its training environment to an unseen new one, it often performs poorly. These are distinct problems faced by the same class of models; ultimately, we aim to address both simultaneously, improving model performance across different problem scenarios.

42HOW Robotics: The paper’s particularity lies in the fact that you did not modify the original VLA or retrain the policy, but instead adjusted the environment before task execution. What prompted you to choose changing the environment rather than the model?

Prabin Rath: This question has two parts. First, StageCraft does not imply that methods directly modifying the model are bad; those approaches are certainly important as well.

A policy can fail in two ways. The first is that the model itself lacks the capability to control the robot to complete the task. For example, if a task requires dexterous manipulation, such as picking up a difficult-to-grasp object, this is a skill issue. If the model does not possess the required skills, adjusting the environment will not help.

The second type of failure is caused by the environment. In this case, the policy might already have the ability to complete the task; it simply needs slight environmental adjustments to make the task easier. The policy inherently possesses the capability but requires minor auxiliary intervention.

Another reason is that retraining these models is often impractical in many deployment scenarios. Imagine a company handing you a robot equipped with its proprietary VLA; you may have no idea what data it was trained on. When the robot is deployed to a kitchen and encounters an unfamiliar object on the countertop, continuously collecting data and fine-tuning the model is not feasible because fine-tuning takes time. Manual data collection requires significant effort. In contrast to other VLA optimization approaches like reinforcement learning (RL), StageCraft offers much higher sample efficiency: it requires only a few policy rollouts to understand the cause of failures and handle faults. This is because vision-language models (VLMs) possess strong reasoning priors, enabling them to comprehend the reasons behind failures.

Identifying "Disturbances" from Execution Logs

42HOW Robotics: You use past successful and failed execution logs to infer which objects in the current environment might be causing failures. What is the most difficult part of this process? How does the system determine whether a specific object could be the cause of failure?

Prabin Rath: StageCraft follows the logic of Monte Carlo estimation. The VLM, serving as the reasoning backbone, reviews multiple execution records of the VLA. We run the VLA in different environments where the policy might encounter various conditions.

Imagine a robot in your home kitchen attempting to complete a task. It sometimes fails and sometimes succeeds. We provide the outcome labels: when it fails, we tell the system it failed; when it succeeds, we tell the system the task was completed. The VLM’s role is to reason about which objects in the environment might be related to the failure.

This reasoning occurs within the context of the VLM. After accumulating sufficient execution samples, the algorithm in the prompt decomposes the environment into different sets of objects, then searches for the minimal set of objects associated with the failure.

For example, if the policy consistently fails when a water bottle appears in front of the camera, the VLM may infer that the water bottle is correlated with a higher probability of failure. It might then suggest removing the water bottle or moving it to a position with less interference. This is a hypothesis made by the VLM to improve the VLA's performance.

42HOW Robotics: You do not remove all potential disturbances from the environment, only those likely to cause failure. Is this primarily to preserve the original task structure, or to avoid introducing new failures through intervention?

Prabin Rath: The second reason is the main one; we want to avoid introducing new failures.

Interventions may affect the objects involved in the task itself. For example, in the cup-stacking demonstration shown, the cups are the task objects. The more changes made to the environment, the greater the likelihood of introducing new failure modes.

Therefore, we only adjust the environment when the model is highly confident that a specific obstacle or disturbance is likely to cause VLA failure. The model still makes estimates, but adopts conservative judgment to avoid performing actions beyond what is necessary. We aim to apply only the minimal intervention required to help VLA succeed.

The paper employs two distinct VLAs: Pi0.5 and SmolVLA. SmolVLA is relatively weaker, while Pi0.5 is stronger and can operate with more distractors. When the number of distractors increases to five, six, or seven, the workspace becomes extremely cluttered; only then do we see Pi0.5 begin to fail. If it could already handle one or two distractors, why would we need to alter the environment? In the future, if VLAs become stronger and are no longer affected by distractors, StageCraft should adjust accordingly by making no environmental changes at all. It should be able to determine that the underlying VLA has become sufficiently capable.

**Applicable Scenarios and Future Directions

42HOW Robotics: The mainstream approach to enhancing VLA capabilities today involves scaling up the model, collecting more robot data, and performing post-training. You delegate certain problems to environmental adjustments for resolution. Based on your research, what types of problems are best suited for this method?

Prabin Rath: StageCraft is specifically designed to handle failures originating from the environment itself, with another starting point being that this method requires no training.

It is particularly suited for uncommon scenarios. For example, if a robot at home suddenly encounters a bouquet of flowers—something you might not have every day and perhaps only receive on special occasions—the robot's performance may begin to degrade when it sees the flowers, even though it functions well in its usual home environment. To the robot, the flowers are out-of-distribution items that it has never seen before.

You do not want a robot to perform poorly simply because an out-of-distribution object appears. A more generalizable VLM can help the VLA understand that this new object might lead to failure and make simple adjustments, allowing the policy to maintain good performance. Our starting point is to avoid failures that could have been prevented.

42HOW Robotics: Suppose future VLAs can understand the current environment, predict their own risk of failure, identify which objects might cause failure, and move them away before the task begins. At that point, will your mechanism still need to exist as an independent system, or can it eventually become a foundational capability integrated into a general-purpose VLA?

Prabin Rath: This is a promising research direction. We certainly hope models will become smarter in the future; unfortunately, they have not yet reached that level. But even if models become smarter, we may still need stronger mechanisms to supervise how they operate.

StageCraft currently proposes a method for handling VLA failures. In the future, the question may shift to VLA safety: even with more capable models, are they safe enough? At that stage, supervision mechanisms might focus on safety or on personalized intervention—what matters to one person may not matter to another.

Therefore, we can reframe the problem. The overall goal is for higher-level supervisory mechanisms to help action-centric models like VLAs handle task details, so they do not have to bear all high-level reasoning themselves. High-level reasoning can be delegated to models that are stronger and better suited to handle such problems.