TacVLA Team: When Should Tactile Sensing Be Used More Cleverly?
"Who has this idea? It's how you integrate it."
"At IROS 2026, when discussing the addition of tactile sensing to VLA models, Zhang Heng was less concerned with whether to use touch and more focused on a specific question: When exactly should robots employ tactile feedback?"
"Equipping robots with tactile sensors and feeding those signals alongside visual and language data into models is no longer a novel concept."
"The real challenge lies in the fact that within a single manipulation trajectory, tactile information does not hold equal value throughout."
"During free-space motion, before the robot makes contact with an object, tactile signals show minimal meaningful variation; once contact occurs, visual changes may be slight, but contact forces begin to shift rapidly."
"This is precisely the problem addressed by the paper accepted at IROS 2026, 'TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation.'"
"The paper features co-first authors Zhang Kai Di and Zhang Heng, supervised by Arash Ajoudani from Yi Da Li Ji Shu Yan Jiu Yuan iit and She Yu from Pu Du Da Xue."
"Zhang Heng completed his doctoral research this summer and is currently conducting postdoctoral studies at Carnegie Mellon University (CMU). During his PhD, he focused on learning-based methods for contact-rich manipulation under the supervision of Arash Ajoudani. He also conducted visiting research in She Yu's team at Pu Du Da Xue, which led to the collaboration with co-first author and PhD student Zhang Kai Di on this TacVLA paper."
He had long focused on whether robots could autonomously determine, based on the current task and state, when to be compliant versus when to apply force, how much force to apply, and how to safely interact with the environment.
With the rapid development of VLA models, this question has been brought into a new model framework.
Zhang Heng recalls that in the very early stages, when he deployed related models locally, he found that vision and language enabled the model to understand scenes and tasks increasingly well, but many states remained difficult to stabilize from visual input alone, especially once true physical contact occurred. To address this, he introduced CompliantVLA-adaptor, which adds a stiffness adjustment mechanism to VLA actions, allowing raw actions to modulate compliance according to task states.
"Vision alone is often insufficient."
In Zhang Heng's view, this is one of the reasons why tactile sensing should be integrated into robot learning systems as soon as possible.
Consequently, he visited the MARS lab under the guidance of Professor She Yu at Pu Du Da Xue (Purdue University) and completed this work alongside Zhang Kai Di. Zhang Heng noted that the MARS lab has consistently conducted leading-edge research in the field of tactile sensing; the group presented 10 papers at the recent IROS conference.
In the TacVLA paper, the authors explore whether the model can engage tactile input only when it provides genuinely useful information, rather than constantly. This approach aims to broaden and enhance VLA performance while minimizing the computational burden imposed by additional modalities.
Gating Mechanism: Adding a "Switch" to Tactile Input
TacVLA employs a tactile-array sensor featuring a 15×8 grid to map the distribution of contact across its surface.
Instead of encoding the tactile map into a high-dimensional, image-like representation, TacVLA uses a lightweight MLP encoder to compress the 120 readings into 36 tactile tokens, which are then fed into the VLA alongside language and visual data.

The rationale is straightforward: treating tactile data as dense imagery would rapidly extend the token sequence due to the volume of pixels, thereby increasing the computational burden on the Transformer.
This point resonates with the discussion on tactile representation in CONTACT:CONtact-aware TACTile Learning for Robotic Disassembly. The two works address different problems, but both deal with the same practical constraint: more tactile information does not mean all of it is worth including in the policy.
After resolving tactile representation, TacVLA posed a further question: when should these tactile tokens engage with the model? To answer this, the team designed a contact-aware gating mechanism.
The system first determines contact using a fixed threshold; the contact flag switches from 0 to 1 only when the number of taxels exceeding a preset pressure threshold surpasses a specific count.
This is a binary judgment. When no contact is detected, the model masks the tactile tokens via attention and zeros out their corresponding embeddings; once contact occurs, these tokens rejoin cross-modal interaction.
However, this gate controls only tactile tokens. Visual, language, and other information are not switched on or off together by the gating mechanism. TacVLA does not dynamically decide whether vision or touch is more important at a given moment; instead, it makes a more specific decision about whether tactile information should be allowed to enter the model.
TacVLA currently implements a strategy where tactile tokens are masked when there is no contact, and are permitted to participate in cross-attention only after contact is detected.
Having Touch Does Not Mean Knowing How to Use It
The team designed two sets of real-robot experiments.
The first category consists of four constraint-locking disassembly tasks: removing an object from a tight-fit shaft, extracting an object after pressing a latch, pulling out an object after twisting it 90 degrees, and sliding an object inward before pulling it out. After completing the disassembly, the robot must place the object into a bowl. These tasks require extensive fine-grained manipulation because objects that look similar cannot be distinguished by vision alone.

In other words, approaching and grasping the object is only part of the operation; the robot must also perform actions to release constraints in order to successfully extract the parts.
To distinguish the respective roles of tactile input and the gating mechanism, the team compared three configurations: a fine-tuned π0.5 without tactile input, a version with tactile input but without gating, and the complete TacVLA.
The average success rates across the four tasks were: fine-tuned π0.5 at 63.75%; with tactile input but without contact-aware gating at 71.25%; and the full TacVLA at 83.75%.
Looking only at the averages, one might draw a simple conclusion: adding tactile data helps, and adding gating improves it further.
However, when broken down by individual task, the results are less uniform. On Task 3, vision-only fine-tuned π0.5 achieved a 65% success rate; adding tactile input without gating actually dropped this to 60%; the full TacVLA improved it to 70%.
This indicates that simply adding tactile data does not guarantee better performance. Moreover, each configuration was tested only 20 times per task, meaning the difference between 60% and 65% on Task 3 corresponds to just one successful trial (12 vs. 13 successes).
The team also observed failure modes. Without gating, the robot might misalign when approaching objects, repeatedly attempt to re-grasp, get stuck in intermediate states, or lift failed attempts. The paper speculates that tactile tokens fused unconditionally may interfere with visual localization before stable contact is established.
The second category is bin-picking. The front-facing camera cannot see inside the box, and the wrist-mounted camera faces occlusion and limited lighting once inside. The robot must locate, grasp, and extract objects under visually constrained conditions.

In this task, TacVLA ultimately achieved a 70% success rate, while fine-tuned π0.5 reached only 10%; the two diffusion-policy baselines scored 0% and 5%, respectively.
Both diffusion-policy baselines incorporate tactile input and are not vision-only models. The distinction lies in their training methodology: they are trained from scratch using task-specific data, whereas TacVLA is built upon a pretrained π0.5 backbone.
The paper attributes the significant performance gap primarily to the priors provided by pre-training, rather than simply crediting it to "TacVLA having tactile capabilities."
Beyond standard task execution, the team tested performance under visual impairment by occluding the front-facing camera.
Under occlusion, fine-tuned π0.5 achieved an average success rate of approximately 30%, while TacVLA reached 62.5%—roughly 2.1 times higher. However, tactile sensing did not replace vision; TacVLA’s performance still dropped from 83.75% (unoccluded) to 62.5%.
The team also demonstrated tests involving human interference. When the robot had already grasped an object and begun moving it out of a box, a person suddenly pushed the object back inside. Instead of mechanically replaying its original trajectory, TacVLA re-explored the environment, re-established contact, and re-grasped the object. Fine-tuned π0.5 failed to recover.
Across these experiments, TacVLA demonstrates that when vision is insufficient to fully describe the physical state, tactile feedback enables the policy to adjust actions based on the current state; and this information yields significantly better results when introduced only upon actual contact, rather than being applied throughout the entire trajectory.
This aligns with the paper’s emphasis on contact-aware tactile fusion. The abstract describes this design as selectively activating tactile tokens only upon detecting contact, thereby avoiding irrelevant tactile interference while allowing visual, linguistic, and haptic inputs to jointly enter the Transformer.
Beyond Fusion: Considering Feedback Latency
Once the experiments were successfully run, the questions did not end there.
Zhang Heng believes that a core challenge facing tactile VLAs is that different modalities do not operate on the same information and time scales.
Vision provides scene understanding and higher-level planning cues; once contact is established, changes in force and touch may require the robot to correct its actions more promptly. He described this difference using the concept of "two loops" in the interview.
"After contact, these changes in force and touch... it's essentially two loops with different frequencies."
Therefore, regarding how touch/force might eventually integrate into larger foundation-model-style systems, Zhang Heng distinguished between two possible architectures.
One approach involves having a relatively high-level VLA handle semantic understanding and planning, while adding a higher-frequency tactile/force feedback layer downstream.
The other follows the direction currently being explored by TacVLA: directly fusing touch into the model, where it participates alongside vision and language in the attention layer to generate actions.
It remains difficult to determine which architecture will become the final standard.
This question has also come up repeatedly in other tactile-focused interviews at 42HOW Robotics during IROS.
Zhang Zhi Yuan discussed a layered approach separating high-level visual reasoning from higher-frequency contact refinement, while Adeesh Desai, co-first author of CONTACT, favored having slower high-level models exchange information with faster low-level controllers.
It is clear that as tactile input begins to enter foundation-model-style policies, the research question has moved beyond "whether to use touch" to: What representation, time scale, and model layer should tactile signals operate at?
After Tactile "Speaks Up"
Turning back to TacVLA itself, its experimental boundaries are clearly defined. The study fine-tuned and evaluated the model on four decomposed tasks and one bin-picking task.
The paper lists three limitations. First, gating relies on binary threshold rules, preventing continuous, learned adjustment of modality weights; second, the spatial resolution of the tactile array is limited, restricting inference on fine contact geometry; third, evaluation focused on short-horizon, contact-heavy tasks, leaving extension to longer-horizon, more complex tasks for future work.
These boundaries prompt further thought: What gaps must be filled to move from effective methods in a few tasks to a tactile VLA covering hundreds or thousands of tasks?
Zhang Heng did not simply attribute the challenge to a lack of tactile data.
"We cannot talk about touch in isolation," he said. TacVLA is built on broader VLA research; the current work uses π0.5, and with better foundation models in the future, tactile research must continue to explore how to integrate with them. Foundation model capabilities, tactile representation, training methods, and multimodal fusion all have room for improvement and are mutually reinforcing.
On the other hand, the data foundation for tactile signals is indeed significantly weak. Zhang Heng noted that many existing robot datasets lack force/tactile signals. Historically, the vast majority of data used to train VLAs has been based on vision and robot state, meaning tactile learning faces a fundamentally different data landscape from the start.
Even if data volumes increase substantially in the future, there is no simple answer to what constitutes truly 'good' tactile data.
"Perfect data is not necessarily the best."
Zhang Heng even suggested that some data that appears imperfect or noisy may not be bad data at all. He cited an example of moving an object from point A to point B. Human intuition might favor smooth, clean motion trajectories; however, less smooth, seemingly 'wobbly' processes can introduce greater data variation, helping the model achieve better robustness and generalization.
For the tactile VLA paradigm to truly scale, a series of foundational questions must be answered.
In Zhang Heng's view, VLAs have already pushed robots toward stronger general-purpose generalization. Looking ahead, dexterity, contact, and safety will become increasingly unavoidable challenges.
This was also his direct takeaway from this year's IROS.
When discussing the vision of robots entering homes, Zhang Heng remains cautious about timelines, suggesting it may still be many years away. However, he is certain that when a robot picks up an object like a cup or a strawberry, it cannot simply "know what to do"; it must also understand how to make contact and apply the appropriate force to complete the operation.
TacVLA provides a concrete starting point, but this step naturally raises questions about feedback frequency, the scale of tactile data, and how different modalities should ultimately be coordinated architecturally.
For Zhang Heng, the next step is clear: "From a research perspective, we have always believed that the next stage is dexterous manipulation and touch."
Beyond Research
Regarding his career plans, Zhang Heng notes that seeking an academic position is one path, though it represents a choice influenced by precedent. At the same time, he is actively engaging with investors to explore entrepreneurial opportunities. He feels that historical trends and his professional expertise have converged, and given his connections over the years, he should consider how to create greater value for society.
He did not disclose specific details regarding the startup direction, but confirmed it will be in the field of physical AI. In the long term, his ultimate goal is to pursue the role of a Robot Scientist. He aims to build tens of thousands of robot scientists operating 24/7, leveraging underlying AI agent systems combined with real-world robotic execution experiments to contribute to a major explosion in human knowledge discovery.
A few years ago, he engaged in deep discussions on this topic with Zhang Peng Song, a doctoral student at Duo Lun Duo Da Xue. These insights were detailed in their co-authored position paper, Scaling Laws in Scientific Discovery with AI and Robot Scientists.
Speaking enthusiastically, he said, "Humanity spent over four hundred years establishing the modern scientific research paradigm, driving rapid advancements in physics, chemistry, biology, medicine, and other fields. It is precisely because of this knowledge and discovery that we have today's civilization and well-being. We believe that the true era of scientific discovery driven by AI and robotics—the Age of Autonomous Discovery—has just begun."
In the future, robots and AI will usher in an unprecedented acceleration of scientific discovery. When asked about his dreams for the future, he smiled and said that while robotics is one aspect, he actually hopes to become a film director and win an Oscar.
Let’s wait and see.
