CONTACT Authors: What Tactile Representation Does Robotics Need?

Over the past few years, the focus of robotics research has been shifting.

As sensing, learning, and foundation models become increasingly integrated, new questions are emerging: Which information perceived by a robot actually helps it execute actions effectively?

Tactile sensing is at this inflection point. For a considerable period, research on tactile sensing focused primarily on sensor design and enhancing perception capabilities; now, with advances in robot learning, tactile data is increasingly being applied to dexterous manipulation, robot learning, Vision-Language-Action (VLA) models, and world models.

However, simply adding a new sensing modality does not guarantee improved robotic manipulation capabilities.

The IROS 2026 main conference paper "CONTACT: CONtact-aware TACTile Learning for Robotic Disassembly" addresses this question directly. The paper was authored by Yosuke Saka, Jyun-Chi Hu, Adeesh Desai, Zhang Zhi Yuan, and others, with She Yu as the corresponding author.

Rather than merely discussing whether robots need tactile sensing, CONTACT focuses on two key questions: First, which manipulation tasks truly benefit from tactile sensing? Second, for policy learning, which tactile representations are more effective?

During IROS, 42HOW Robotics spoke with CONTACT authors Zhang Zhi Yuan and Adeesh Desai. The two authors shared their perspectives on contact-rich manipulation, tactile representation, and how multimodal policies can effectively leverage tactile information.

Zhang Zhi Yuan is currently conducting research on robot learning at the MARS Lab of Pu Du Da Xue, with a focus on 3D policy learning, vision-tactile multimodal learning, and world models for contact-rich manipulation. Early in his career, he primarily focused on vision-based tactile sensing and sensor design, and participated in the development of GelRoller. After joining Pu Du Da Xue, he gradually shifted his attention from "how to obtain tactile information" to "how robots can truly use this information to complete operations after acquiring it."

Adeesh Desai, who participated in research at Pu Du Da Xue MARS Lab, is a co-first author of the CONTACT paper. His research focuses on tactile learning, contact-rich manipulation, and physical interaction. In CONTACT, he primarily explored how different tactile representations affect policy learning, and how structured contact information helps robots perform complex operations.

Their research approach also aligns with the core question CONTACT aims to answer: robots do not need more sensing, but rather task-relevant physical information. This is what distinguishes CONTACT from many studies that simply add sensor modalities. Instead of continuing to ask, "Do robots need more sensors?" it poses a more fundamental question: For a specific task, what exactly is the physical information that the policy truly needs to understand?

The value of tactile sensing is highly task-dependent

CONTACT chose disassembly as its primary research scenario not by chance.

Many existing tactile manipulation works focus on insertion, such as USB insertion or peg-in-hole. These tasks are important, but contact-rich manipulation involves more than just geometric alignment. During disassembly, robots must also deal with friction, shear, deformation, clip release, and changes in local contact states.

For example, a part may be unable to be extracted due to friction; a barb may not yet have disengaged; a latch may need to be pressed down first; and an elastic structure may also have to deform to a certain state before it can be released. The paper summarizes such processes as force-dependent state transitions: task success depends not only on geometric alignment but also on how contact forces and physical constraints evolve with motion.

To this end, CONTACT designed ten disassembly tasks of varying complexity, including five simulation tasks and five real-world tasks. The tasks progress from relatively simple rigid extraction to more complex contact and deformation processes involving lids, barbs, push tabs, and vertical clips.

This set of tasks enables researchers to compare how the role of different sensory representations changes as tasks shift from primarily relying on geometry to increasingly depending on contact state.

Experimental results indicate that the benefits provided by tactile sensing in the tasks tested by CONTACT exhibit clear task dependency. For example, in real-world task R1, the success rates for Vision Only, Vision+TacRGB, and Vision+TacFF were 80%, 90%, and 95%, respectively; in R2, they were 55%, 30%, and 70%; and in the more complex R3, they were 15%, 45%, and 55%.

These results do not imply that 'the harder the task, the more important tactile sensing becomes.' More accurately, when a task relies primarily on geometric alignment, vision often provides sufficient state information; however, when a robot needs to assess friction, shear, deformation, or determine whether a structure is 'stuck' or 'released,' tactile sensing offers physical state information that vision cannot directly observe.

An intuitive example from the paper is centered grasp versus tilted grasp. Due to low contrast and self-occlusion, these two grasping states may appear very similar in external RGB camera views, but tactile signals reveal distinctly different local contact patterns. TacRGB reflects varying surface deformation, while TacFF reveals a more asymmetric shear distribution under tilted grasp.

Thus, vision and tactile sensing are not simple substitutes. Vision excels at describing external geometry and scene state, whereas tactile sensing provides local physical interaction information after contact occurs.

Why Structured Force Field Is More Effective in CONTACT?

Another core set of experiments in CONTACT focuses on tactile representation.

The team uses GelSight to capture tactile information and processes it into two forms. One is TacRGB, a relatively raw, dense tactile image; the other is TacFF, which further represents contact as spatially distributed normal force and shear force, forming a more compact, structured force-field representation.

Across CONTACT’s five simulation tasks and five real-world tasks, Vision+TacFF achieves overall better performance, while TacRGB shows stronger task dependency. For example, in real-world task R2, Vision Only scores 55%, Vision+TacRGB 30%, and Vision+TacFF 70%.

Why does this difference occur?

A key factor is that different tactile representations preserve different information. TacRGB retains rich appearance, surface deformation, and texture cues; if the task goal is material recognition or texture classification, such high-dimensional information can be highly valuable. However, for the disassembly tasks focused on by CONTACT, the policy needs to understand contact dynamics more—such as contact direction, local forces, shear, and whether a structure has been released.

From this perspective, TacFF can be understood as a task-relevant inductive bias. Faced with raw tactile images, the model must learn on its own from massive pixel changes which information relates to action generation; the force field explicitly organizes physical quantities more directly related to contact interaction.

But this result does not mean the force field is a universally optimal tactile representation. More accurately, under CONTACT’s contact-rich disassembly tasks, current data scale, and policy architecture, TacFF provides a structured representation better suited to the current task.

There is a more fundamental distinction: the two are not simply different encodings of the same physical information.

In GelSight sensors, shear force is typically obtained by tracking the lateral movement of markers on the gel surface. CONTACT’s TacRGB, however, is a markerless tactile image that primarily records deformation after the gel is pressed; it does not directly contain such marker motion.

Therefore, under CONTACT’s setup, the difference between TacRGB and TacFF stems not only from representation format but also from differences in observable information itself: some shear information is inherently difficult to obtain from markerless TacRGB inputs, whereas TacFF explicitly provides structured representations of normal force and shear force.

This means one cannot simply assume that scaling up TacRGB data will inevitably allow the model to recover information equivalent to TacFF. The final outcome also depends on sensor observability, data diversity, model capacity, and training objectives. If certain types of physical information do not enter the observation space, adding more of the same type of data cannot fill this gap out of thin air.

Of course, raw tactile images are not without value. TacRGB retains rich appearance, surface deformation, and texture cues, which may still play important roles in material recognition, texture perception, or large-scale learning scenarios. How different tactile representations might be combined as data scale and model capabilities improve remains an open question.

More modalities do not automatically mean better policies

If TacRGB and TacFF contain different information, would feeding both into the policy simultaneously yield better results?

CONTACT’s experiments offer a notable finding: not necessarily. In the R3 and R5 experiments, simply adding both TacRGB and TacFF did not outperform using TacFF alone; in some cases, performance dropped significantly.

This result first demonstrates a key point: the effective fusion of heterogeneous sensory inputs does not occur naturally simply because there is "more information." Additional observation dimensions can simultaneously introduce redundancy, inputs with differing statistical properties, and more complex representation learning challenges.

However, it is important to distinguish between experimental observations and potential mechanistic explanations. CONTACT directly observes that naive multimodal fusion yields no additional benefit; whether the performance decline stems from input redundancy, encoder capacity limits, fusion architecture choices, or optimization difficulties under limited data requires further investigation.

This naturally points to subsequent directions: truly effective multimodal policies may require modality-aware attention mechanisms, specialized encoders designed for different modalities, or dynamic selection of which sensory information should enter the policy based on the current interaction state.

Therefore, what CONTACT aims to emphasize is not "giving robots fewer sensors," but rather that for manipulation policies, the sheer volume of information is not the sole objective; task relevance and representation quality are equally critical.

From Perception to Action Correction

Beyond success rates, CONTACT also observed behavioral differences across various sensory configurations.

Notably, TacFF does not always enable the robot to complete tasks faster. In some successful rollouts, policies with force feedback undergo longer interaction times. One possible reason is the additional computational overhead introduced by tactile processing; on the other hand, continuously changing force cues provide the policy with further basis for adjusting actions.

This is tied to the characteristics of contact-rich manipulation. Before contact occurs, visual observations change continuously as the robot moves, while tactile signals may remain nearly constant; after contact actually happens, external visual changes can sometimes be small, whereas local tactile signals change rapidly.

Dim-light experiments further demonstrate the complementary relationship between vision and touch. In R5, Vision Only performance dropped from 15% under normal lighting to 0% in dim light, while Vision+TacFF maintained 55% under both conditions. However, in R1, low illumination still caused a significant decline across all configurations.

Therefore, touch is not a universal 'blind-spot filler' for when vision fails. More accurately, the two modalities describe different layers of state: vision provides global geometry and scene context, while touch provides contact-local physical interaction information.

CONTACT ultimately proposes not adding more sensors, but re-examining how robot policies should select, represent, and utilize physical information.

The key to touch scaling lies not just in data volume

As robot learning moves toward large-scale data and foundation models, touch faces a challenge less prominent in the vision field: real-world tactile data is difficult to obtain directly from the internet.

The internet contains massive amounts of RGB images and videos, but these data typically lack synchronized contact force, shear, deformation, or tactile state information. Consequently, tactile learning cannot simply replicate the scaling path that vision models rely on web-scale data.

Zhang Zhi Yuan is currently focusing on a direction called tactile imagination: using real paired vision-tactile data to learn the relationship between vision and touch, then enabling the model to predict task-relevant tactile representations from visual observations. If this cross-modal prediction is sufficiently reliable, it could potentially transform data that originally contained only vision into representations more useful for physical interaction.

Adeesh Desai further emphasized that predicted tactile data and actual measured tactile data are not necessarily in a substitute relationship.

In TacImag, for contact-sensitive tasks, the imagined tactile force fields generated by the model have already demonstrated effects close to those of real tactile sensors. However, visual perception still struggles with hidden or sudden situations: such as occluded contacts, friction, the onset of slippage, or whether a latch has truly popped open.

Therefore, he prefers to treat the tactile model's predictions as prior information, with sensor measurements serving as correction terms. If there is a significant discrepancy between the predicted tactile state and the actual measurement, this discrepancy itself may become an important signal: it indicates that a physical interaction occurred in the real world that the model did not anticipate.

From this perspective, as tactile prediction capabilities improve, the role of real tactile sensors may also change: they may no longer always serve as the primary source of information, but instead play more of a role in identifying model prediction errors and correcting physical understanding.

Another promising avenue to explore is egocentric human manipulation data. Humans can naturally perform a large number of dexterous interactions, but how to extract physical interaction information from first-person vision, hand movements, and limited tactile measurements for transfer to robots remains to be explored.

In this light, the questions raised by CONTACT do not stop at "whether robots need tactile sensing."

As tactile learning continues to scale, more critical questions may be: what kind of physical information is worth perceiving? What kind of representation can preserve interaction states relevant to the task? When and in what manner should this information enter the policy?

For contact-rich manipulation, what truly needs to scale may not just be the volume of data, but also the robot's ability to represent, predict, and utilize physical interactions.

References:

CONTACT : https://vict0rhu.github.io/CONTACT-Website/

TacImag:https://tacimag.github.io/