Interview with Shallow-π Authors: Compressing Transformer Layers Improves Embodied AI Performance

A cylinder and a rotating perforated plate are enough to create a performance gap between two robot models.

In Shallow-π’s real-robot experiments, the task required a robotic arm to pick up a cylinder and insert it into a moving hole. A larger model with more layers might miss the target while waiting for computation results; a smaller student model with fewer layers, whose output is close to that of the teacher, can adjust its actions faster based on new visual input.

This work, authored by Boseong Jeon, Yunho Choi, and Taehan Kim, and affiliated with San Xing Yan Jiu Yuan, proposes a knowledge distillation method for flow-based Vision-Language-Action (VLA) models. It simultaneously reduces the Transformer depth of both the vision-language backbone and the action-generation head from 18 layers to 6 layers, enabling a smaller 'student' to learn the action-generation capabilities of a larger 'teacher.'

On standard operation benchmarks, Shallow-π achieves more than a two-fold inference speedup, with an absolute success-rate drop of less than one percentage point. On the ALOHA platform and Jetson Orin, end-to-end real-robot deployment latency drops from 364 milliseconds to 110 milliseconds; in 10 tests of the pin-in-hole task, the teacher completed 7 attempts while the student completed all 10. For the RB-Y1 equipped with a dexterous hand, the team also validated tasks such as garbage sorting, lid opening, and cylinder placement on Jetson Thor.

These results raise a question: How should the capabilities of robot models actually scale? As environments continuously change, there remains a gap between action quality in offline evaluations and task performance on real robots, separated by computational latency.

Around this issue, some researchers believe that completing architectures, increasing parameter counts, and training with more data through Scale Up may lead to better generalization. Others constrain the relationship between the 'brain' and the 'body' into compact spatial representations by focusing on different model paradigms and benchmarks.

The Boseong Jeon team appears to take a different path—they compressed the Transformer layers of the π model. On standard operation benchmarks, Shallow-π increased inference speed by more than twice, with an absolute success-rate decline of less than 1 percentage point, achieving state-of-the-art performance among compressed VLA models.

At the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 42HOW Robotics interviewed Boseong Jeon, the first author of the study, to discuss the technical choices behind Shallow-π.

Boseong Jeon: First author of Shallow-π. He received a Bachelor’s degree in Mechanical and Aerospace Engineering from Shou Er Da Xue in 2017 and a Ph.D. in Robotics in 2022 under the supervision of H. Jin Kim. He subsequently joined San Xing Yan Jiu Yuan to conduct research on world models. His research interests cover robot motion planning, edge-side generative AI, robot foundation models, and world models, with a focus on accelerating Vision-Language-Action (VLA) models through knowledge distillation to facilitate their deployment on edge devices and real-world robots.

Yunho Choi: Holds a Bachelor’s degree in Electrical and Computer Engineering from Shou Er Da Xue and earned his Ph.D. from the same institution in 2024 under the supervision of Songhwai Oh. He previously served as an AI Robotics Researcher at Sequor Robotics and is currently a Staff Engineer in the Robot Intelligence Team at San Xing Yan Jiu Yuan. His research focuses on vision-based robot learning, particularly decision-making and reinforcement learning for embodied agents, including Vision-Language-Action (VLA) models for humanoid robots.

Taehan Kim: Holds a Bachelor’s degree in Mechanical Engineering from Yan Shi Da Xue. He worked as a Robotics Research Engineer at San Xing Yan Jiu Yuan from 2022 to 2026 and joined Ludo Robotics in August 2026 as an AI and Machine Learning Research Engineer. His research focuses on learning-based robot manipulation and humanoid robots, covering VLA models, imitation learning, and reinforcement learning fine-tuning for dexterous manipulation, dual-arm assembly, and humanoid control. He is currently working on building foundation models that support social interaction for humanoid robots.

The Key Is Latency

42HOW Robotics: Could you start by introducing your academic background? Why did you choose to research VLA models later?

Boseong Jeon: My initial background was in traditional robotics, not artificial intelligence or machine learning. My doctoral work primarily focused on traditional motion planning methods, but that field has changed significantly since then.

While at San Xing Yan Jiu Yuan, I began researching generative models such as diffusion models and participated in applying AI image removal features to Galaxy smartphones. I wanted to combine these two areas of expertise to research learning-based motion planning. I also aimed to make these models run faster on devices to support broader applications. That is why I chose to research VLA models.

42HOW Robotics: Specifically regarding Shallow-π, what led you to focus your research on the issue of 'models being too slow'?

Boseong Jeon: I believe that when deploying humanoid robots or other robots on edge devices, a primary bottleneck is inference speed. The models are often too slow to achieve real-time operation.

42HOW Robotics: How did you determine that latency was indeed the problem?

Boseong Jeon: We need to consider the update frequency of camera feeds. For example, if the system receives 10 frames of color images per second, but the motion planning runs at a much lower frequency, the planning cannot reflect the latest observations. Consequently, actions are based on older frames, which naturally degrades performance.

Additionally, if computation is slow, the robot's movements themselves appear unnatural: it moves once, waits for the next action to be calculated, and then moves again. It exhibits a stop-and-go behavior and lacks fluidity. This is why I aimed to reduce inference latency.

42HOW Robotics: On the Jetson Orin, Shallow-π₀ achieves an end-to-end computation time of 110 milliseconds, compared to 364 milliseconds for the teacher π₀. Task success rates improved from 7/10 to 10/10 for pin-in-hole insertion, from 6/10 to 9/10 for apple scooping, and from 12/20 to 17/20 for garbage sorting, while computation time dropped from 130 milliseconds to 78 milliseconds. In these real-world tasks, the student model outperformed the teacher. To what extent does this improvement come from distilled action quality versus faster processing of observation data?

Boseong Jeon: This question lies at the heart of the key message we wanted to convey in the paper. If the student model’s outputs are close to the teacher’s while running significantly faster, it may actually perform better on real robots during deployment.

In this dynamic scene, a turntable is rotating. The robot on the left uses the teacher model, while the one on the right uses the student model. The teacher takes longer to compute actions and relies on earlier observations.

Because it is slower, it cannot update its observations in time to inform the action, so it fails to insert the cylinder. In contrast, the student can update its actions based on newer observations more quickly, resulting in better performance. This is the main point I want to make.

After executing a portion of the actions, we request a new set of actions. If the new trajectory differs significantly from the previous segment, smoothing between them may cause issues. When new action blocks arrive, we need to smoothly connect them with the previous ones to ensure motion continuity. However, if the difference between the two segments is too large, smoothing may actually degrade the result.

The effect is better when the two segments are closer together, and faster inference makes this scenario more likely. Therefore, as demonstrated, I believe that using the student model improves overall operational accuracy.

"Simplifying Complexity"

42HOW Robotics: We know that VLMs handle images and language, while the action head uses this information to generate robot actions. The π model's Transformer architecture passes conditional information between layers. Reducing the number of layers could affect model performance, yet you chose to compress both the vision-language backbone and the action generation head from 18 layers to 6. This is a critical decision. Why compress both parts simultaneously? What were your considerations?

Boseong Jeon: We need to reduce the computational load of vision-language models. VLA models like π, which perform well, feature a tightly coupled architecture where the vision-language model and the policy head are connected at every layer. Because of these interconnections, removing layers from only one part feels awkward.

We aim to remove layers from both components simultaneously rather than just one. Some other algorithms prune only the upper part of the backbone, but those layers may contain valuable information; deleting them could result in significant information loss.

These connections themselves are also important. When compressing both parts synchronously, performance might improve if this connection structure is preserved. Therefore, I wanted to maintain this information-passing structure.

Once the connection structure is preserved, the shallow model needs to learn how to utilize it. Thus, the paper’s training scheme consists of three parts: supervised learning based on real actions, having the student match the teacher’s predicted velocities, and aligning the attention of action tokens with the intermediate layers of visual-language information.

The attention alignment applies only to action tokens and is placed in the intermediate layers. For a 6-layer model, the success rate was 93.0% using only task loss; it rose to 93.9% after adding teacher output supervision; and further increased to 94.6% with the addition of attention distillation. These results provide further evidence for the design choice of 'preserving information passing'.

42HOW Robotics: Why did you adopt uniform layer selection when initializing the student model?

Boseong Jeon: Because it is simple and effective. I have also tried other layer initialization methods, such as those based on feature similarity or proposed in other papers. Those approaches offer elegant explanations, but ultimately, with enough training steps, the results converge to similar outcomes.

Using heuristic rules to select initial layers only complicates matters and hinders usability.

42HOW Robotics: For what usage is it unfavorable?

Boseong Jeon: Model usability. I prefer simple designs. Having worked in industry, I avoid overly complex methods. I want this project to be accessible regardless of the user’s background. That is my design philosophy.

42HOW Robotics: Does this imply that the original model contains many redundant layers?

Boseong Jeon: Yes. In the paper, I demonstrate how success rates drop when specific layers are skipped. You can see that skipping certain layers causes no decline in success rate at all, indicating redundancy. I found that inter-layer similarity varies with denoising time, and the degree of similarity does not reliably reflect functional importance. Even if we prioritize removing the layer with the lowest single-layer sensitivity, task success rates still plummet as more layers are removed. Shallow-π’s approach is to uniformly sample layers to initialize the student model and then recover capability through training; simply cutting 18 layers down to 6 does not yield the same result.

Other curves illustrate how features change across different Transformer layers. In early layers, feature changes are significant; by the middle layers, these changes diminish. Both observations point to the existence of redundancy.

Why is there redundancy? My guess is that the model is first pre-trained on a very large dataset, but then only used for a small set of tasks. Therefore, some of its capabilities are not fully utilized, which makes sense.

Of course, 'simple' does not mean skipping training. When the number of training steps is sufficient, we select initialization layers based on layer sensitivity, but this did not yield additional benefits. So, after removing the complexity of selecting layers that showed no extra benefit, we shifted our focus to distillation.

42HOW Robotics: You trained teacher and student models for ALOHA and RB-Y1 respectively. For those hoping to unify models across different robot platforms, can this distillation method also be applied?

Boseong Jeon: I think this distillation method is relatively easy to apply to such cases. Newer models are attempting to use a single model to unify different robot platforms. Our distillation method does not impose restrictions on this scenario; I believe it can be extended to those models.

42HOW Robotics: If you continue to reduce the number of layers, visual tokens, action generation steps, and lower numerical precision, combining these compression techniques, which capability do you think would be most affected first?

Boseong Jeon: Visual understanding.

When deployed on small edge devices with limited computing power, the student model's final action accuracy or closed-loop stability might actually be better. But for small models, visual understanding becomes more difficult.

When discussing the trade-offs of compression, Boseong Jeon redirected the question toward whether distillation can preserve the diverse choices present in a teacher model. He explained using a distribution with two possible outcomes.

Boseong Jeon: Suppose the teacher model’s distribution has two modes. During distillation, one possibility is that the student selects only the larger mode. In this case, the student fails to cover the other mode. This represents a limitation of distillation.

In our operational scenarios, the range of options is limited. Compared to image or video generation, distillation may be more effective. Images and videos have numerous possible outcomes; for example, if asked to generate an image of a cat, there could be millions of variations in its appearance.

However, if the task is for a gripper to hold a phone, it simply needs to grasp the phone. It does not need to employ this specific grip, that grip, or any other alternative method.

Training costs do not disappear either. Because distillation requires loading both the teacher and student models simultaneously, the training overhead exceeds that of direct layer skipping. In future research, we may selectively freeze certain components, filter for training samples with higher information content, and combine visual token compression with diffusion step reduction to further improve inference efficiency.

Moving Toward the Real World

42HOW Robotics: As large model capabilities improve, will robots evolve into hierarchical systems where large models handle planning while small models manage real-time control?

Boseong Jeon: This discussion has been ongoing for several months. The core challenge lies in managing the complexity of deploying robots in the real world: should we rely more on vision-language models (VLMs) and large language models (LLMs), or on VLA architectures themselves?

When using tool calling, tools should remain lightweight. On the other hand, large models—such as world-action models or VLAs—also need to run on edge devices. For device-side applications, knowledge distillation is equally necessary.

After entering industrial applications, I believe models should run on devices rather than relying on the cloud. Therefore, regardless of which path is taken, knowledge distillation will be necessary in the future.

42HOW Robotics: How do you view the future application prospects of this research?

Boseong Jeon: It is a natural workflow to first develop a large model and then distill it into a smaller model for customers with limited computational resources. I believe the robotics industry can offer products in both forms.

42HOW Robotics: Do you mean that robots could have two versions—one based on a large model and another on a small model?

Boseong Jeon: Exactly. Take the iPhone, for example: there is the Pro version and the Plus version. The same logic applies here. In my view, this is an inevitable path toward commercialization: start with a large model, then perform lightweight fine-tuning for specific skill subsets to adapt to particular application scenarios.

42HOW Robotics: From our perspective, this work offers valuable insights for practical deployment. In China, the industry is increasingly focused on how to land these technologies in real-world scenarios, rather than just focusing on the models themselves.

Boseong Jeon: As you mentioned, real-world deployment and simulation benchmarks are fundamentally different. In simulation, the environment step can wait for action generation; the environment does not need to advance until an action arrives. But in the real world, the robot must continue executing actions even before new ones are generated.

This means it operates in an open-loop execution state, continuing based on older observations, which can degrade performance. Therefore, in real-world applications, we need to generate the next set of actions faster. This is critical for practical use.