Conversations with the IROS Best Paper Award Finalist Team: Unveiling the Secret Behind Robots' 'Glass Dizziness'
From September 27 to October 1, the 39th IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026) was held at the David L. Lawrence Convention Center in Pittsburgh, Pennsylvania.

The conference is expected to draw more than 4,000 attendees, with nearly 2,000 papers presented and 145 institutions exhibiting. This marks the first time Pittsburgh has hosted IROS in 31 years. The conference chair is Howie Choset, a professor at the Robotics Institute of Carnegie Mellon University (CMU).

The most notable highlight is the list of finalists for the IROS 2026 Best Paper Award. The conference received 4,348 submissions, setting a new record high. Of these, 1,585 were accepted, resulting in an acceptance rate of approximately 36%, the lowest in recent years. Only ten papers were nominated for the Best Paper or Best Student Paper Awards.

Among them is work from the team led by Academician Zhang Hong of Southern University of Science and Technology, focusing on glass surface reconstruction to enable safe robot navigation. The core contribution is a training-free framework that uses depth foundation models as structural priors. It employs a robust local alignment method based on RANSAC (Random Sample Consensus) to fuse with raw sensor depth data.
This approach naturally mitigates interference from erroneous glass measurements, restoring accurate metric scales. Additionally, the newly designed GlassRecon RGB-D dataset is specifically tailored for robot navigation, allowing ground truth for glass regions to be derived through geometric deduction. The technology holds significant application potential, enabling robots to clearly identify glass surfaces or objects made of glass and navigate around them during daily operations.
The study’s core contribution lies in addressing the limitations of existing sensors. Glass surfaces disrupt indoor robot navigation, causing depth sensor measurements to become severely distorted. While foundation models like Depth Anything 3 provide excellent geometric priors, they lack absolute metric scale. Co-first author Yu Jing Wen believes that the research was likely selected for the IROS 2026 Best Paper/Best Student Paper Award because it cleverly combines cutting-edge depth perception foundation models to solve robot glass avoidance problems, with its conceptual simplicity earning recognition from the program committee.
But why is recognizing glass planes so critical for robot navigation? During IROS 2026, 42HOW Robotics interviewed Yu Jing Wen, co-first author of the shortlisted team, to uncover the intricate relationship between robots and glass.

Interviewee Profile: Yu Jing Wen is a Ph.D. student at the Zheng Jia Chun Ji Qi Ren Yan Jiu Yuan (CKSRI) of Xiang Gang Ke Ji Da Xue and a member of the Shen Zhen Shi Ji Qi Ren Shi Jue Yu Dao Hang Zhong Dian Shi Yan Shi at Southern University of Science and Technology (SUSTech). He completed his undergraduate studies in the Dian Zi Yu Dian Qi Gong Cheng Xi at SUSTech, advised by Chair Professor Zhang Hong and Professor Tan Ping. His main research interests include visual navigation, 3D reconstruction, and SLAM (Simultaneous Localization and Mapping).
Starting Point
42HOW Robotics: Could you briefly introduce your academic journey?
Yu Jing Wen: I entered Southern University of Science and Technology in 2017, joined the Dian Zi Yu Dian Qi Gong Cheng Xi at Southern University of Science and Technology in 2019, and came into contact with robotics in early 2019. In the summer of 2019, during an exchange at National University of Singapore, I encountered a project related to autonomous mobile robots and found myself interested in this direction.
Coincidentally, in 2020, two faculty members in the Department of Electronic and Electrical Engineering at Southern University of Science and Technology focused on robotics; one of them was my later supervisor, Chair Professor Zhang Hong. Professor Zhang worked in Canada for 30 years and is a member of the Canadian Academy of Engineering. Since 2020, I have been conducting research on mobile robots under his guidance. After graduating with my bachelor’s degree in 2021, I continued working with him for over a year, gaining foundational experience in mobile robotics. With his recommendation, I entered Xiang Gang Ke Ji Da Xue for my Ph.D., where I am advised by Professor Tan Ping, an expert in 3D vision.

During my PhD, Professor Zhang Hong was my co-supervisor, and I also conducted research training at the Shen Zhen Shi Ji Qi Ren Shi Jue Yu Dao Hang Zhong Dian Shi Yan Shi at Southern University of Science and Technology. Our lab had many members; I was fortunate to collaborate with one of Professor Zhang’s master’s students, Zheng Jia Min, who is the first author of this paper. Together, we completed the work that was selected as a Best Paper Finalist at IROS this year.
42HOW Robotics: What sparked this research?
Yu Jing Wen: Around late 2024, Zheng Jia Min contacted me, hoping I would join her on a project. Another author of the paper, Chen Guang Cheng, was also part of the team; he had been working on reconstructing objects from reflective surfaces.
42HOW Robotics: Does “reflective” refer specifically to glass reflections?
Yu Jing Wen: Objects on reflective surfaces are not quite the same as glass. For example, ceramic surfaces are also reflective, regardless of their color. Chen Guang Cheng used a specialized sensor—a polarization camera—to perform reconstruction. Glass is a particularly special material in daily life; light passing through it undergoes refraction or reflection. During our discussions, we gradually thought that perhaps we could do something specific with glass.
With this opportunity, we began working in this direction. It also evolved from previous projects and has practical significance. This brings us to Mei Tuan.
42HOW Robotics: You did indeed thank Mei Tuan Yan Jiu Yuan in your acknowledgments. What is the relationship between your team and Mei Tuan?
Yu Jing Wen: We have been collaborating with the autonomous delivery division of Mei Tuan for approximately three years. They utilize compact robots to deliver food in environments such as airports and large shopping malls. This project has already been deployed at Shenzhen Bao'an International Airport.

After operating on flat ground within a single floor, they aim to enable these robots to navigate stairs, similar to hotel service robots. However, unlike hotels, airports do not grant easy access to elevator control systems for every robot. Therefore, the robots must be capable of autonomously entering elevators and performing necessary operations.
During this process, we identified a specific challenge: many elevator interiors feature reflective metal surfaces. In such environments, depth cameras provide inaccurate perception data. Starting from Mei Tuan's practical problem, we determined that addressing glass—a relatively direct and simpler scenario—would be an effective first step, which ultimately led to this research topic.
42HOW Robotics: Has this research outcome been deployed yet?
Yu Jing Wen: The method described in the paper has not yet been deployed. We have conducted some related deployment research previously.
42HOW Robotics: The fact that glass obstructs robots is counterintuitive. From a human perspective, we assume that identifying glass and glass-like objects is not difficult. What kind of obstacles does glass pose to embodied AI? Do traditional sensors fail to see it entirely, or do they misidentify it?
Yu Jing Wen: Our current depth sensors fall into two categories: depth cameras and LiDAR. Both measure distance using light. However, glass surfaces are unique; light is either reflected by them or passes directly through. Consequently, these sensors become ineffective when encountering glass.
We want to know the depth of glass surfaces, i.e., their geometric shape in 3D space, whether it is a mirror, a glass door, or a glass cup on a table. Currently, doing this with existing sensors is quite difficult.
This leads to two types of errors. One is holes, where we don't know there is an object; the other is misestimating the position of the surface. For example, when glass is very transparent, sensors might think the surface is further back. In short, its depth perception is inaccurate.
This brings about two categories of problems. One is the focus of our paper: navigation, where robots may crash into glass. The other is tabletop manipulation, especially in embodied AI scenarios, where robots cannot accurately locate objects, leading to operational errors.
42HOW Robotics: Among all materials, are glass surfaces and glass objects the most difficult for sensors to handle? Some other metal materials may also have reflection or refraction, but is glass the most prominent among them?
Yu Jing Wen: You could say that. This is also related to some reflective properties inherent to the material itself. The characteristic of glass is its high transmittance; light passes through easily. In contrast, metals always reflect light, but after reflection, it becomes diffuse, or the intensity becomes unpredictable.
Additionally, glass is relatively common in our daily lives and in robot application scenarios. Especially when we want robots to enter homes or offices, it is more common than special metal materials. Of course, in industrial scenarios, metal is also an important issue that needs to be addressed.
Approaching the Problem Simply
42HOW Robotics: Compared to previous research, such as CDM, SILICA, and GlassGuard, what do you think is the biggest innovation of this paper? What made it shortlisted for the IROS Best Paper/Best Student Paper Award?
Yu Jing Wen: We compared many mainstream data-driven methods. What we actually wanted to do was fill in the gaps in depth sensor measurements. In this regard, the LingBot-Depth approach developed by the Antlingbo Robotics team is very close to the problem we are trying to solve.

Being selected as a finalist is partly due to luck, but also because our paper’s approach remains simple and direct. The method proposed in the paper is quite straightforward, with ideas that are unpretentious and clear—there are no flashy elements or major theoretical innovations, but the results are strong. We mainly tried applying state-of-the-art models to areas where robots need them most. The IROS program committee likely values how, given the rapid development of large models and AI, these technologies can be practically deployed in robotic applications, especially in realistic scenarios.
There are now many advanced monocular depth perception models. We organically integrated them into the entire robot navigation system. As long as a model can output scale-invariant monocular depth estimation maps, our framework can adapt to such models.

42HOW Robotics: The paper proposes a training-free modular pipeline, particularly featuring local RANSAC. Could you briefly explain its novelty? Did you try any approaches that ultimately failed?
Yu Jing Wen: We tried many approaches. First, we had to define the problem clearly. We have a depth map generated by Depth Anything V2 or Depth Anything 3, which estimates the depth of glass surfaces. However, the issue with such depth maps is that they lack true scale; what they estimate is relative depth.
Such depth maps cannot be used directly for robot reconstruction. For example, they might output depth values ranging from 0 to 255, but this does not reflect actual distances from 0 meters to 255 meters. Our sensors can obtain true-scale data, but the challenge is that they fail when encountering glass. However, in real-world scenarios, there will always be larger areas not covered by glass, which serves as an assumption of our method. So we asked: How can we leverage these unaffected, correct sensor depths to make estimations?

We adopted a straightforward approach: first detect the glass regions, then use depth information from non-glass areas to recover scale. While this works, experiments showed that training such a model requires substantial data, and glass types vary significantly across different scenes.
Another idea was for the team to train a small network to estimate glass regions, but this ultimately remained limited by data availability.
42HOW Robotics: Where exactly did the data issues lie?
Yu Jing Wen: The difficulty lies in annotating glass. Our annotation relied on a planar assumption: since the depth of a glass region is consistent with its edges, we could define it using more than three points. This process required significant manual effort, primarily handled by Zheng Jia Min.
We also explored using VLMs for pre-annotation (pseudo-labeling) to accelerate the workflow. We initially annotated approximately 1,000 images intended for training, but midway through realized that the high precision required for this type of data made the manual cost prohibitive. We had previously attempted data-driven methods, but the results were unsatisfactory.
We then considered sampling based on patch partitioning, using a training-free method to identify non-glass regions. Specifically, we used RANSAC, assuming that non-glass areas constitute the majority of the scene.
42HOW Robotics: The paper divides the RGB-D dataset into Easy and Hard versions. How does this assist with glass recognition?
Yu Jing Wen: You can simply understand it as the proportion of the glass plane within the entire image. Areas where the glass occupies a small range are included in the Easy dataset, while those with a large glass presence go into the Hard dataset. This data helps sensors reconstruct and recognize glass planes.
"Ghost Glass"
42HOW Robotics: How does your final 3D dense mapping guide robots to plan motion trajectories in specific scenarios?

Yu Jing Wen: The paper lists several scenarios, particularly those involving glass doors. If the map lacks depth information at the glass locations, the robot assumes the area is passable and crashes into it. We have video evidence of this, though we did not include the demo in the paper. In one real-robot experiment, which was quite amusing, the robot repeatedly crashed into the same spot. Therefore, current capabilities are primarily focused on static obstacle avoidance.

42HOW Robotics: Your experimental results and demonstrations show significant improvements over previous studies in depth testing. Sensors can navigate around glass doors in real-world scenarios, but you also noted some failures. What limitations has the team observed in the application of Glass Recon?
Yu Jing Wen: Our failure cases occur when the glass area is relatively large. Our method works effectively when the glass region occupies a smaller portion of the image compared to non-glass regions.
42HOW Robotics: If the robot enters an environment like a mirror maze, does it fail completely?
Yu Jing Wen: Yes, it fails completely. Glass Recon is built upon depth priors—the article title includes “via Depth Prior.” This depth prior itself is a neural network. When that component fails, our method cannot work either.
42HOW Robotics: Are there other limitations you haven’t mentioned?
Yu Jing Wen: Another limitation is that we still rely heavily on a pre-trained model capable of estimating relative scale depth. We are attempting to deploy a large model onto actual robots and have designed a specific approach for this purpose. If future advancements improve such models, our method will benefit accordingly. However, if those improvements encounter unresolved edge cases, our method would also be unable to handle them.
The Team’s Next Steps
42HOW Robotics: You mentioned that the team might try integrating models like SAM 3 or GPT to optimize glass obstacle avoidance and recognition. Is this part of your next plan?
Yu Jing Wen: Yes. We are also working on new projects, such as attempting to integrate SAM 3. SAM 3 is a general-purpose segmentation large model. Or more simply, use GPT directly: give it a prompt saying "segment the glass area in this image," and then apply our method. This is what we are currently doing. Later, we will likely try to deploy our solution on panoramic cameras to perform indoor panoramic visual navigation.
42HOW Robotics: Is there potential for this technology to be integrated into more physical robots for testing in the future?
Yu Jing Wen: I already have a company in Shenzhen, where I am co-founding a startup with classmates. In the future, for specific scenarios, we will both refine this method and conduct more real-robot experiments. Our company primarily targets office environments, where glass doors and large glass panels are common; we will certainly continue to apply the methods proposed in our research.
Additionally, I believe the embodied AI industry is gradually cooling down. People are realizing that purely end-to-end approaches are not feasible. At this stage, it is better to return to classic robot architectures and explore how they can support further advancements in embodied AI, enabling true on-site deployment.
42HOW Robotics: Embodied AI currently has many technical routes. Compared to current approaches, what specifically does your mentioned classic robot architecture emphasize?
Yu Jing Wen: Simply put, an entire robot system can be divided into perception, planning, and control layers. Under a classic robot framework, there is no cross-embodiment problem.
For example, although our work and demos are performed on robots, many mapping results were captured using smartphones. Any device equipped with an RGB-D camera can use our method; no specific robot configuration affects the algorithm. In a classic framework, planning is also embodiment-agnostic; control involves different parameters for different robots. Thus, the cross-embodiment issue found in VLA paradigms does not exist. As the focus shifts toward generalization across different embodiments, the classic robot framework may become more suitable.
42HOW Robotics: Embodied AI technologies have slowly come together like grains of sand forming a tower. Many researchers like you are increasingly integrating these technologies into models and robots. What do you think the future of embodied AI looks like?
Yu Jing Wen: Over the past one or two years, research based on VLA paradigms and world model paradigms has emerged. Seeing the success of large language models, everyone wants to replicate similar success on robots. Our work offers an insight: a robot’s working environment is far more complex than simple visual question answering, chatting with large models, or writing. Subtle changes in lighting can alter the information perceived by the robot’s sensors at any given moment.
In this context, I believe relying solely on a single end-to-end model makes it difficult to complete such a series of complex robotic tasks. Following our company’s technical route for entrepreneurship, we will also layer the entire robot system according to the classic robot framework. Each layer may utilize the most advanced models, leveraging LLMs and other models to handle specific problems.
Through years of research and deployment in robotic systems, we identified several specific challenges. We then discovered that state-of-the-art models could address these issues. Leveraging our experience in robotic system integration, we organically combined these solutions. The result is a hybrid approach that requires no network connectivity; it relies purely on computation, enabling rapid deployment on small-scale robots.
