Touch Is Becoming the Next Data Modality for Robot Training | IROS 2026

On September 27, the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) opened in Pittsburgh, USA. 42HOW Robotics was on site.

Several session rooms were completely full, with indoor seats taken early and doors kept open to accommodate standing attendees outside. While IROS covers a broad range of topics across all robotics directions, the sessions that drew the largest crowds on the first day were exclusively dedicated to tactile sensing.

42HOW Robotics attended three of these sessions, featuring different teams and research directions.

Professor Edward Adelson from the Massachusetts Institute of Technology presented on vision-tactile research, starting with the classic GelSight technology and systematically outlining the technical routes and engineering trade-offs of various tactile solutions.

Assistant Professor Li Yun Zhu from Columbia University introduced work related to FlexiTac, explaining how low-cost tactile sensors can be integrated into robot learning systems to adapt to data collection for both simulation training and real-world scenarios.

Qi Hao Zhi from Amazon Frontier AI & Robotics used dexterous manipulation as an entry point, focusing the latter half of the presentation on the deployment strategies for the OSMO tactile glove and ForceBand devices.

The three presentations offered different perspectives, but all pointed to a long-standing technical gap in the industry: contact physics information beyond vision is difficult to collect and model at scale.

As robot learning evolves and data scale becomes central to research, a practical contradiction has emerged: during interactions with the real world, robots undergo extensive contact processes, yet many critical physical states have long gone unrecorded, unmodeled, and unused.

This is the core reason tactile research has seen concentrated growth this year.

Contact That Vision Cannot Capture

Tactile sensing is not a recent research direction.

In 2009, Micah Johnson, then a doctoral student at Massachusetts Institute of Technology (MIT), and MIT professor Edward Adelson published early work at CVPR that laid the foundation for subsequent GelSight technology. Their study used a structure combining a transparent elastomer with a reflective film to fully capture minute textures and surface deformations when an object presses against it. A camera at the bottom then captured these deformation details using multi-directional lighting.

This approach converted contact physics information, which traditional sensors struggle to quantify directly, into image signals that computer vision can process directly.

For over a decade, GelSight has gradually become one of the mainstream technical approaches in robotic tactile sensing, with the industry successively developing various other sensing solutions such as piezoresistive, capacitive, piezoelectric, and magnetic types. Although the iteration of tactile technology has never ceased, it has never achieved the status of a standardized basic configuration for robots when compared to vision.

As summarized in the 2025 IEEE Chu Jue Ji Qi Ren Zong Shu (IEEE Tactile Robotics Review), current mainstream technologies still cover multiple parallel routes including optical, piezoresistive, piezoelectric, capacitive, and magnetic methods. The structural forms, installation positions, and output signal standards of sensors are highly dispersed, and the industry has yet to form a unified paradigm.

Adelson pointed out this status quo at IROS: today, the industry hardly discusses whether robots need vision, as cameras have long become standard equipment; however, tactile sensing remains an optional configuration for most robots rather than a mandatory capability.

The shift in industry perception occurred during the development of foundational models for robotics. For the past two years, the industry has generally driven robot learning by focusing on data scale. To address the scarcity of real-robot data, teleoperation data, internet videos, and first-person human demonstration data have been extensively reused.

Vision can record scene appearance at low cost, and trajectories can record hand movement paths, but a long-overlooked shortcoming has been exposed: vision can record what the world looks like and where the hand moves, but it cannot record what happens when the hand contacts an object.

The grape-grabbing experiment presented by Li Yun Zhu at IROS intuitively confirmed this issue. When the robot grabbed grapes from a bag, the bag created visual occlusion, preventing the wrist-mounted camera from capturing the true interaction state between the fingers. Strategies trained solely on vision could output standard grasping actions but were prone to empty grabs or crushing objects by gripping too tightly.

With tactile feedback integrated, robots can obtain real-time contact status between their fingertips and objects. Even if subjected to external disturbances during experiments, the robot can autonomously adjust its posture to re-complete the grasping task whenever it detects an abnormal gripping state or failure to stabilize against the object, thereby ensuring interaction stability.

Grape grasping is merely a typical example; the same logic applies to handling transparent test tubes. Vision systems struggle to precisely determine minute positional and angular deviations of a held test tube. In delicate tasks such as insertion, rotation, wiping, and tightening, physical interaction states—including slippage, friction, normal force, and shear forces—remain entirely within the blind spot of visual perception. When humans grasp objects, they continuously monitor grip stability, sliding trends, and applied forces through touch, dynamically adjusting their actions. Purely vision-based robots lack this core closed-loop perceptual capability.

This means that tactile data is far more than simple 'additional force control data.' It encompasses rich physical information, including contact location, pressure distribution, multi-dimensional force vectors, slippage states, local geometric deformation, surface texture, and material properties.

Different sensing technologies capture only partial dimensional features, resulting in vastly different data formats: GelSight outputs high-resolution surface deformation features via visual imaging; FlexiTac collects pressure distribution through piezoresistive arrays; OSMO achieves three-axis contact force measurement by relying on magnetic field variations; and ForceBand breaks away from fingertip-sensing logic, estimating fingertip applied forces through forearm electromyographic signals.

This complexity is also why tactile modeling is significantly more difficult than visual modeling.

The camera industry has established a mature, unified standardization system, allowing image data from different devices to be uniformly mapped into a two-dimensional pixel space for adaptation to various visual models. Tactile sensing lacks such a unified paradigm. Switching sensor solutions fundamentally alters the corresponding physical observation dimensions, data representation, and signal logic. This inconsistency is a key pain point hindering the generalization and scaling of tactile technology.

Making Tactile Sensors Affordable

At this year's IROS conference, Li Yun Zhu used his team's self-developed FlexiTac sensor as the core focus to outline the primary solutions for scaling tactile sensing and reviewed the team's technological iteration over many years.

She began with a classic study from her PhD days at MIT. In 2019, Subramanian Sundaram and Li Yun Zhu’s team published in Nature a scalable tactile glove featuring 548 piezoresistive sensing units integrated onto a knitted substrate, with the entire array costing only about USD 10.

The team wore the glove to manipulate 26 types of objects, collecting 135,000 frames of full-hand tactile data. Using convolutional neural networks, they achieved object recognition and weight estimation while dissecting the collaborative interaction patterns across different regions of the human hand. The paper’s core insight remains highly relevant in today’s industry context: robotic vision has rapidly adapted to machine learning systems largely due to massive datasets and mature collection paradigms; at that time, however, the tactile domain lacked deployable large-scale sensing platforms and corresponding datasets.

By 2019, the industry had already begun exploring solutions to increase tactile data volume, but the scale of over ten thousand frames remained vastly inferior to the millions of hours of video and millions of motion trajectories relied upon by current robot foundation models.

More critically, the cost of collecting tactile data cannot be naturally reduced through hardware storage expansion. Visual data collection is passive; cameras can continuously shoot to generate vast amounts of video. Tactile data, however, depends on real physical interaction, requiring repeated pressing, sliding, and friction to generate valid signals. This process not only continuously wears out sensor hardware but is also subject to multiple variables, including robot structure, sensor morphology, installation position, and material degradation.

This is why Adelson emphasized sensor durability as a key topic at this conference. To obtain high-precision geometric details, visual-style tactile sensors require extremely thin flexible surface layers. The thinner the film, the more accurately it reproduces deformation details, but the risks of wear, tearing, and delamination increase significantly, creating an inherent trade-off between resolution and durability. He cited an example: using 2 mm thick high-toughness rubber (similar to inner tubes for truck tires) can greatly enhance sensor durability, but spatial resolution drops directly to the millimeter level, failing to capture subtle interaction features.

Experiments demonstrating durability close to practical applications were also presented at the event. Retired chemical engineer Dick Cottrell used abrasive-tipped indenters to strike sensors under a constant force of 4 kg, accelerating device wear and aging. Test results showed that mainstream GelSight Mini and DIGIT sensors failed after approximately 30 minutes, whereas modified versions equipped with TPU protective layers remained stable for up to six days. Adelson acknowledged that the industry currently lacks unified standards for testing tactile sensor durability, leaving durability assessments without normative guidelines.

For small-scale individual studies, replacing damaged sensors is manageable. However, once entering large-scale data collection phases, issues of hardware consistency are magnified infinitely: after replacing a new sensor, are its output signals compatible with old data? Do historical datasets become invalid? Can training strategies adapted for robots still be reused after changing the sensing “skin”?

AnySkin, introduced in 2024, directly addresses this pain point. This magnetic tactile sensor features a modular split design that separates the easily worn sensing surface from the underlying electronic hardware, enabling rapid skin replacement while optimizing signal consistency across different components. The team conducted rigorous cross-device generalization experiments: data was collected and policy trained using a single sensing skin; after replacing it with a new skin without any recalibration, the original model was reused directly.

In high-frequency contact tasks such as plugging in connectors, swiping cards, and inserting USB drives, AnySkin’s policy performance dropped by only 15.6% after a skin change, with the best test scenarios showing a decline of just 13%. In contrast, ReSkin suffered a performance decay of up to 43% under the same conditions when its sensors were replaced, highlighting the significant advantage of the modular design.

Li Yun Zhu presented FlexiTac, outlining a low-cost pathway suited for large-scale deployment. Relying on mature piezoresistive principles, it perceives pressure distribution through an array layout. It does not pursue the ultra-high geometric resolution of GelSight; instead, the spacing between sensing units is controlled at around 2 mm. This scale is close to the human fingertip’s two-point discrimination threshold of 2–3 mm. Although human and machine perception mechanisms cannot be fully equated, this design effectively meets the sensory needs of daily fine-interaction tasks.

FlexiTac’s core breakthroughs lie in manufacturing efficiency and cost control.

The first-generation product required 30 minutes of manual preparation per unit, costing approximately USD 5. After process optimization in the second generation, the production time for a single unit was compressed to 3 minutes. Cost advantages become even more pronounced at scale: mass production of 20 units costs about USD 4.19 each, while bulk orders of 1,000 units drop the price to USD 1.36 each. The sensor’s hardware and software are fully open-source, supporting high-frequency acquisition at 100 Hz. It can be bent or cut freely and adapted for installation on grippers, fingertips, and full-body robot skins. Its official positioning is clear: low cost, open-source, scalable, and tailored for large-batch data collection needs.

A simple reduction in cost is not the core value; a USD 1.36 sensor will not directly boost robot performance. The true industry innovation lies in the reconstruction of the evaluation system for tactile sensors. The focus has shifted from solely targeting single performance metrics like accuracy and resolution to adding core dimensions suited for data infrastructure, including manufacturing efficiency, batch consistency, hardware lifespan, cross-device data compatibility, and simulation reproducibility.

During his presentation, Li Yun Zhu shared long-term stability test cases: four early FlexiTac sensors made for NVIDIA in the summer of 2024 remained fully functional after a year and a half. A policy model trained on data collected one year ago remains stable and effective today. To address the common challenge of batch inconsistency in soft-material sensors, the team completed specialized optimizations. They also simplified the physical perception mechanism by modeling sensor characteristics using a spring-damper system in simulations. With only a few parameters, the simulated tactile signals can closely match real-world conditions, effectively reducing the randomization costs associated with simulation training.

The industry’s logic for evaluating tactile sensors has broadened significantly. Early research focused primarily on “how accurately it measures and how finely it distinguishes.” Today, greater emphasis is placed on “whether it can be mass-produced, whether parts can be reused, cross-device universality, simulation adaptability, and long-term data validity.”

The key milestone marking the formal integration of tactile sensing into robotic learning systems is its deep binding with existing mature data acquisition solutions. The widely adopted UMI acquisition paradigm relies on handheld grippers to perform operations in real-world environments, using vision and positioning systems to record demonstration trajectories; this approach offers greater flexibility compared to fixed-robot data collection. Native UMI devices are equipped with mechanical springs that provide human operators with force feedback during operation, but they cannot retain tactile data, resulting in the permanent loss of a vast amount of fine-grained interaction information.

With FlexiTac integrated, the same demonstration workflow can simultaneously capture visual and tactile multimodal data, supporting both model pre-training and policy fine-tuning alongside teleoperation data. This seemingly simple hardware upgrade significantly enriches data dimensions: whereas past demonstration data could only record "visual frames and executed actions," it can now fully reconstruct contact, force, and slippage details throughout the interaction process.

Tactile signals also exhibit strong scene sparsity and high value: during the robotic arm's free movement, the tactile channel provides almost no valid information; however, upon contact interaction, tactile sensors instantly output high-temporal-resolution, highly precise physical signals that support rapid dynamic response. The recently popular tactile world models have specifically adapted to this characteristic. For example, FeelWorld initiates deep vision-tactile joint modeling only after contact is triggered, distinguishing between the modal requirements of free motion versus dense interaction, thereby further improving model training efficiency and inference accuracy.

Supplementing Touch from Human Data

If future data relies primarily on autonomous collection by robots, low-cost, durable, and replaceable tactile sensors could address part of the problem. Meanwhile, this round of robot foundation models has another important data source: humans.

Internet video, first-person footage, motion capture, and wearable devices all leverage an intuitive fact: humans perform a vast number of operations daily that robots still struggle to replicate. Collecting human operation samples directly is far less costly than the alternative approach of training robots to complete tasks first and then relying on those robots to generate similar data.

Qi Hao Zhi introduced OSMO during the latter half of the IROS session, targeting this specific gap in human demonstration data.

Video can capture the entire process of a person unscrewing a bottle cap, wiping a table, squeezing a sponge, or holding a glass. With tracking algorithms, it is possible to obtain hand posture and motion trajectories, but standard video cannot record contact force information. OSMO equips its robotic hands with tactile gloves that feature 12 triaxial magnetic tactile sensors arranged on the fingertips and palm. When soft magnetic materials are compressed and deformed, underlying sensors read changes in the magnetic field to calculate interaction information such as normal force and shear force. A notable design approach for OSMO is that the same set of gloves can be used for both human data collection and robot hardware deployment, narrowing the gap between human and robot data at the sensor interface level.

In wiping task experiments, the policy was trained entirely on human demonstrations without introducing real-robot data. With the addition of tactile information, the task success rate reached 72%, outperforming the pure vision baseline. The reduction in failure cases was primarily concentrated on issues related to contact pressure.

The value of this experiment lies not just in the 72% figure, but in raising an increasingly critical question: If human demonstration data is to serve as training material for robot policies, what specific information from human operations must we actually record? Video data is easy to scale up; smartphones, cameras, and the internet have already built mature infrastructure, and hand trajectory tracking solutions are becoming increasingly refined. However, there are no existing massive data sources for information such as force, slip, and contact position, requiring additional hardware setups for data collection.

Relying solely on tactile gloves for data collection makes it difficult to scale indefinitely.

Qi Hao Zhi pointed out two practical pain points on site: wearing tactile gloves for continuous work of 12 hours is not user-friendly, and sensors suffer continuous wear and tear from repeated contact. Consequently, the team’s subsequent ForceBand adopted a different technical path. Shaped like a wristwatch, it is worn on the wrist and integrates an eight-channel surface electromyography module with an inertial measurement unit to record forearm muscle activity. A model then converts these signals into estimated forces for each finger.

During the training phase of the EMG2Force model, subjects still needed to wear tactile sensors on their fingertips. These fingertip sensors provided ground truth, while sEMG and inertial signals served as model inputs. Once the model learned the mapping relationship between the signals, the fingertip sensors could be removed during formal human demonstration data collection, leaving only the wristband and camera.

Related work on ForceBand released in June this year came with a 10-hour multimodal dataset, including first-person-view video, sEMG, inertial data, and fingertip force information. After completing short-term individual calibration for new users, ForceBand can automatically reconstruct the force variation curve for each finger in demonstration samples. Team test results showed that compared to pure vision-based force estimation, this solution reduced force prediction error by more than 50%. In grasping tasks that require controlling grip force for different objects, the task success rate reached 87%.

This approach still fails to fully resolve all challenges in tactile data acquisition.

Myoelectric signals can only indirectly estimate fingertip force, exhibit individual differences among users requiring separate calibration, and show limited capability in reconstructing absolute force values.

Qi Hao Zhi explicitly noted at IROS that the current solution is better suited for capturing relative changes in force. However, the progression from OSMO to ForceBand reveals a clear research shift: researchers are beginning to decompose the costs of multimodal data collection and evaluate them item by item.

If collecting human operation data at the scale of millions of hours becomes feasible in the future, considerations must extend beyond the number of data collectors to include the trade-offs of various sensing hardware worn on the head, hands, and wrists. Cameras record scene appearance, tracking modules capture motion trajectories, while myoelectric devices or tactile gloves supplement contact and force information. Adding each new perceptual modality introduces overheads in wearability comfort, individual calibration, temporal synchronization, hardware lifespan, and data cleaning.

For datasets of equal duration, the density of physical interaction information they carry may vary drastically.

General Representations of Touch and World Models

Declining hardware costs and established data collection pipelines solve only the first half of tactile research.

To truly integrate touch into foundation models, another core question must be answered: Can massive amounts of tactile data from different sensors, objects, and tasks learn a general representation?

Computer vision has ImageNet, with billions of internet images and a stable, mature pixel representation system. Tactile data volumes are much smaller, and there is significant variation between sensors. Even when looking only at visual tactile sensors, differences in light sources, gel layers, marker points, and camera angles across devices create noticeable discrepancies.

In 2024, teams from Meta and Hua Sheng Dun Da Xue launched Sparsh, advancing research along the path of general tactile representation.

The team employed self-supervised learning to pretrain a tactile encoder using a dataset of over 460,000 tactile images covering multiple visual tactile sensors. They also built TacBench, an evaluation benchmark that tests model capabilities such as identifying material properties and completing operational planning across six task categories.

Experimental results showed that representations obtained through self-supervised pretraining improved average performance on TacBench by 95.1% compared to end-to-end training tailored for single tasks or single sensors. This approach mirrors the development trajectory of computer vision: first pretraining backbone networks with large amounts of unlabeled data, then adapting them to downstream tasks using small amounts of task-specific data.

However, tactile foundation models remain in their early stages.

In June this year, RCT research from TU Dresden re-examined the true level of generalization capability in tactile models.

Using three DIGIT sensors, the team repeatedly pressed against 122 industrial reference materials, collecting 29,279 frames of tactile data while fully preserving the entire process of each press.

The study revealed an easily overlooked pitfall: consecutive tactile frames collected during a single press exhibit high similarity. If the training and test sets are divided randomly by frame, adjacent frames generated from the same physical interaction may end up in both sets. The model appears to generalize well, but this is essentially just recognizing highly similar contact images.

After removing sample overlap between contact sequences, Recall@1 for the touch-to-text task dropped by 17.7 percentage points; when further requiring that test materials be entirely unseen in the training set, Recall@1 fell to just 25.1% ± 6.1%.

This finding reveals a fundamental difference between tactile data and standard images. When a robot finger presses on the same material, capturing dozens of frames from light to deep contact may appear as dozens of independent samples, but physically it represents only a single interaction.

This implies we need to redefine the statistical metrics for tactile data: 460,000 tactile images do not equal 460,000 independent contacts; 1,000 hours of tactile recordings do not equate to 1,000 hours of high-information interactions.

At this year's IROS conference, another trend in tactile research emerged: tactile sensing is beginning to integrate with world models.

Li Yun Zhu presented ongoing work from his lab at the end of his talk, showing a model learning action-conditioned predictions from real interaction data while simultaneously outputting visual and tactile keypoints. In live demonstrations, the model ran at approximately 15 Hz. Before two objects were aligned, predictions captured slippage phenomena; once the test tube and bracket were aligned, visual and tactile predictions jointly drove subsequent insertion actions.

Traditionally, tactile sensors have served as post-hoc perception components: after a robot touches an object, the sensor feeds back the current contact state to the policy module.

World models shift prediction upstream. Before executing an action, the robot can simulate outcomes: where contact will occur, whether slippage will happen, and how forces will change.

Several works have advanced in this direction this year.

TouchWorld employs a layered design for tactile world model prediction, vision-tactile action generation, and high-frequency tactile residual correction. In six long-horizon, high-contact dexterous manipulation tasks, the team measured success rates of 65.0% in clean environments and 53.7% under human disturbance, outperforming corresponding optimal baselines by 15.7 and 18.5 percentage points, respectively. A more recent preprint, DexTouch-WM, reintroduces human demonstration data. This work configures piezoresistive tactile arrays compatible with both humans and robots, redirects human motions to the robot's action space, and trains a world model capable of simultaneously predicting visual and tactile dynamics.

The experiment fixed real supervised robot data at five hours while expanding human tactile data from zero to 100 hours; the pre-training tasks for humans and robots did not overlap. As human data volume increased, the contact Intersection over Union (IoU) on robot test segments rose from 0.415 to 0.588, and the contact F1 score improved from 0.551 to 0.706. Released only in September, this work requires further independent validation, but it transforms the question of whether human tactile data can augment robot world models into a quantifiable problem that can be plotted as scaling curves.

From OSMO’s hundreds of demonstrations and ForceBand’s 10-hour multimodal dataset to DexTouch-WM’s 100 hours of human tactile samples, data scale remains far from internet-image levels, yet the overall direction is clear.

Researchers are guiding touch to replicate the development path of vision: iteratively improving sensor hardware, building datasets, learning general representations from data, and finally embedding those representations into action policies and world models.

The difference lies in the fact that every step forward in touch must additionally bear the constraints imposed by the physical world.