Dialogue with IROS Best Paper Nominee Teams: The 'Anti-Involution' of Dexterous Hands Starts with Moving the Palm

To unscrew a child-proof bottle cap, the most intuitive approach is to make the robotic hand increasingly resemble a human hand.

This typically means adding more fingers, joints, and degrees of freedom, then training a correspondingly complex control strategy.

Zhou Yu Hao from Pu Du Da Xue and his collaborators offer a different answer:

Instead of rushing to add more fingers, let the palm move.

At IROS 2026, Zhou Yu Hao presented the team’s VTAP Gripper. Selected for the Best Paper Award, this design uses only three fingers but transforms the palm—previously often just a passive support surface—into an active contact interface capable of both movement and visual and tactile perception.

The paper is co-authored by Zhou Yu Hao, Sheeraz Athar, Hu Zhi Xian, Huang Bing Hao, Li Yun Zhu, Juan Wachs, and She Yu. Zhou Yu Hao, Sheeraz Athar, Hu Zhi Xian, Juan Wachs, and She Yu are affiliated with Pu Du Da Xue; Huang Bing Hao and Li Yun Zhu are from Columbia University; She Yu serves as the corresponding author.

Lead author Zhou Yu Hao is a fourth-year Ph.D. candidate in Industrial Engineering at Pu Du Da Xue. His MARS Lab, led by She Yu, focuses on integrating mechanism design, tactile sensing, and reinforcement learning. According to the School of Industrial Engineering at Pu Du Da Xue, the MARS Lab and its collaborators have had 10 research papers accepted at this year’s IROS.

In an interview with 42HOW Robotics, Zhou Yu Hao expressed surprise that the paper was shortlisted for the award. The work explores a design space distinct from highly anthropomorphic dexterous hands: Can finger-palm synergy and multimodal perception achieve rich manipulation capabilities using simpler mechanical structures?

How much complexity does a robot need to be truly dexterous? This is the question VTAP aims to answer.

Where Should Degrees of Freedom Be Added?

Current robotic end-effectors generally fall into two extremes.

On one side are single-degree-of-freedom parallel grippers, common in industrial settings. They feature simple structures, stability, and ease of control, but can typically only perform basic grasping and placing tasks.

On the other end are increasingly anthropomorphic dexterous hands. With five fingers and ten to twenty-plus degrees of freedom, they offer a range of motion closer to that of human hands, but this comes with a rapid increase in mechanical and control complexity.

Zhou Yu Hao does not dismiss the value of high degrees of freedom. In an interview, he noted that while higher degrees of freedom enable richer manipulation capabilities, they also require more complex learning-based methods. Furthermore, the mechanical systems are more prone to engineering issues such as overheating and joint failures, resulting in failure rates far higher than those of parallel grippers in factories.

Therefore, their team sought to explore the middle ground between these two extremes: "Would three fingers be sufficient for most tasks?"

According to Zhou Yu Hao, there is no simple linear relationship between control simplicity and dexterity; instead, there may be a more optimal design space.

Therefore, VTAP does not continue to compete on the number of fingers. It has three fingers, each with four degrees of freedom, but the team integrated the palm into the operating system: the palm is no longer just a passive fixed support surface; it can actively approach objects, apply force, stabilize them, and redistribute contact.

The paper summarizes this design as: not adding more contact surfaces, but making existing contact surfaces more useful.

Zhou Yu Hao gave two examples. One is preventing children from opening bottle caps. When humans twist such caps, they often need to press down while rotating. If a robot relies entirely on its fingers to perform these actions, it requires coordinating more joints; on the VTAP Gripper, the active palm handles the downward pressing, while the fingers handle fixation and rotation.

The other example is a syringe. The gripper first holds and adjusts the syringe with two fingers, then directly uses the palm to press the plunger. Fingers and the palm share tasks, reducing reliance on complex finger postures.

He specifically emphasized, 'It doesn't mean it solves new problems, but that using an active palm combined with perception can solve some problems in a simple way.'

This is also VTAP's most core design judgment: dexterity does not necessarily come only from more DoF, but can also come from more effective utilization of existing body structures.

A Single Palm, Two Modes

If the palm is to handle pushing, stabilizing, and holding objects, it must also provide feedback during manipulation.

Getting the palm to move is only the first step. Another core aspect of VTAP is enabling this palm to simultaneously function as a sensing surface.

The team applied a semi-transparent reflective coating to the silicone surface and controls its state via internal LEDs.

With the LEDs off, external light passes through the coating, allowing an internal camera to observe both the area in front of the hand and objects within the hand—effectively acting as an in-hand camera.

When the internal lights are turned on, light reflects off the coating; any surface deformation caused by contact is captured by the camera, switching the system into tactile sensing mode.

Using the same camera and sensing surface, the system toggles between vision and tactile modes with a single power switch that controls the illumination.

In an interview, Zhou Yu Hao further explained that the visibility range and clarity of the visual mode can be adjusted by parameters such as the field of view and surface coating. Once switched to tactile mode, the deformations captured in the images serve as input for subsequent processing tasks, such as geometric reconstruction or force estimation.

The two modes address different information needs during manipulation. Vision observes the relative positions of fingers and objects, while tactile sensing records details of actual contact surfaces.

In demonstrations of autonomous peg-in-hole insertion, the system first uses vision for coarse localization, then switches to tactile mode for finer exploration and geometric estimation. The project page reports that the task was completed with a 1 mm tolerance.

The team also demonstrated autonomous grasping using tactile feedback, as well as syringe adjustment, plunger pressing, and single-object separation within the hand via teleoperation. Single-object separation refers to isolating one object from a group being held. These demonstrations illustrate the roles of perceptual feedback, mechanical capabilities, and teleoperation systems; they are not all accomplished by autonomous learning strategies.

Fingertip tactile sensing and palm perception serve distinct functions. Flexible tactile array sensors are additionally mounted on three fingers. Zhou Yu Hao noted that vision within the palm can observe scenes where fingertips make contact with objects during in-hand manipulation; when the palm also participates in grasping or pushing, its tactile sensing provides additional input. He hopes these modalities will ultimately work together to support manipulation policies.

Changes in contact information are particularly intuitive in demonstrations involving gummy bears and irregular objects for single-object separation. When multiple objects touch the fingers simultaneously, the signal is complex; after separation, the single remaining object in the hand forms a clearer contact pattern.

These sensing methods provide clues for understanding two types of questions: 'where' an object is in the hand, and 'what' is happening between the hand and the object.

How to Control a Three-Finger Gripper with a Five-Finger Human Hand?

However, controlling such a gripper requires overcoming another hurdle.

The human hand has five fingers, while the VTAP has three; their joint arrangements, reachable spaces, and base coordinate directions also differ. These structural differences make direct action transfer between the human hand and the gripper difficult.

Directly mapping human joint angles to the robot struggles to accommodate these geometric and kinematic differences.

Therefore, the team designed a retargeting framework.

First, based on the operator's gesture, it determines which functional grasping category—such as cage, power, or pinch—the action belongs to. This narrows down the range of possible configurations for the robot, reducing ambiguity when one human gesture corresponds to multiple robot poses.

Next, an intermediate coordinate system is introduced to handle the difference between the human wrist coordinate system and the gripper coordinate system. This preserves task-relevant geometric relationships, such as the relative positions and orientations of the fingertips, and adds smoothing constraints to stabilize the motion.

Zhou Yu Hao explained that existing five-finger mappings often use the wrist as the coordinate system, whereas the three-finger gripper operates in 3D space, necessitating a redefinition of the coordinate system. He focused on the positional relationships between fingertips, as well as their distance and direction relative to the hand coordinate system. By solving this in real-time using optimization methods, the system outputs joint control commands for the gripper. He noted that input can be provided via hand skeleton information obtained from a Meta Quest.

With this system, operators can control the three-finger gripper using hand movements for demonstration collection.

VTAP simultaneously records fingertip force, palm vision, and palm tactile data, along with robot state.

Zhou Yu Hao noted that during in-hand manipulation, contact between fingertips and objects, object deformation, and information generated by palm involvement may aid imitation or reinforcement learning.

Thus, VTAP is not just a gripper but also a contact-rich manipulation data collection platform.

He openly acknowledged the system's limitations: the silicone-based optical surface wears down, degrading both tactile and visual signals. He considers this an engineering challenge involving materials; once resolved, it could be deployed in real industrial environments.

Zhou Yu Hao observed that while many embodied AI models already rely on external cameras to understand scenes well, data for fine-grained information requiring tactile perception remains scarce.

One factor limiting the scaling of tactile learning is the past lack of such data and foundation models trained on it. There is no universally accepted best tactile sensor, nor a large-scale dataset comparable to ImageNet. Regarding sharing tactile data across different robots, he admitted he has no answer as to what would serve as the 'common language.'

Currently, the team views VTAP primarily as a hardware and data platform. Zhou Yu Hao stated they aim to use full-hand multimodal sensing to collect more data and further investigate how these modalities should be fused and integrated into policies.

Back to that initial bottle cap: the hand is now moving, and it has richer tactile feedback. But having touch does not mean the robot has learned to use it. In fact, more touch data can sometimes make the policy perform worse.

At IROS 2026, we observed another work from Purdue’s MARS Lab, titled CONTACT: CONtact-aware TACTile Learning for Robotic Disassembly, which asks a deeper question: once a robot can already "feel," what kind of tactile information is actually useful?