Li Xiang Qi Che’s ME-Dex 1.0 Predicts Future Tactile States Alongside Visual Frames
ME-Dex 1.0 enables robots to pre-calculate the contact state that will result from their next action before executing it. The base model team at Li Xiang Qi Che integrated visual, tactile, and motor expert modules into a single denoising framework; here, tactile data is treated as a future observation to be predicted, just like images. By leveraging a hand template, the team aligned tactile information from various dexterous hands and grippers into a unified set of regional features, ultimately improving performance on high-contact-density two-handed tasks.
On September 18, the base model team at Li Xiang Qi Che published the paper "ME-Dex 1.0: Bringing Heterogeneous Tactile Sensing into World Action Modeling." The paper lists ten authors, with Xuancheng Zhang as the first author and Yu Liu as the project lead.
World action models introduce the predictive capabilities of video models into action generation: the model first predicts changes in future frames and then generates actions based on those predictions. Tactile sensing directly reflects physical contact, but historically has mostly served as an auxiliary input connected to the input end of policies or world models. The core idea behind ME-Dex 1.0 is that tactile data and video should both describe the evolution of the world's state, serving as parallel future observations for simultaneous prediction.
This approach requires overcoming three major challenges: the scarcity of tactile data, inconsistent formats across different sensors and robotic hand bodies, and the joint learning of three modalities within a single model. To address these, the team designed an automated data completion platform, a hand template paired with a shared autoencoder, and an architecture featuring shared attention among the three expert modules.
01 Contact Changes Deserve Prediction
During fine manipulation by robots, contact states are typically established gradually in stages. This means actions must dynamically adjust according to the progress of contact; relying solely on visual information is insufficient.
When plugging in a power cord, the hand first touches the edge of the socket, then aligns, and finally presses inward; the force applied differs at each step. Assembling mechanical structures often requires one hand to hold the workpiece steady while the other inserts a pin. When grasping a paper cup, excessive force can crush it, while insufficient force causes it to slip when lifted. These processes are characterized by fingertip forces; images can only capture changes in object contours. Determining whether contact has occurred, if posture is skewed, or if the grip is stable requires tactile signals.
Previous methods for integrating tactile data into policies read the force values at the current moment and used them as conditional inputs, alongside images and language. DECO is a typical example of such an approach: it uses cross-attention to connect tactile features into a pre-trained visual policy, freezing the original parameters and training only the new tactile pathway.
While such approaches enable models to utilize tactile information, they have limitations: the model does not predict the evolution of contact states over subsequent frames. In tasks like grasping and peg-in-hole, forces change dynamically with contact states; the gripper only needs to release once the part is fully seated. What the model truly needs is the trend of how contact states will evolve next, with the current force serving merely as a starting point. The team conducted a controlled experiment by feeding the same set of tactile features into two baseline models, π0.5 and Fast-WAM. Under RoboTwin's Random configuration, adding tactile data improved Fast-WAM's success rate by 1.00 percentage points, whereas it caused π0.5's success rate to drop by 3.08 percentage points. This demonstrates that simply placing tactile data at the model's input yields varying benefits depending on the baseline architecture itself.
The base model team at Li Xiang Qi Che proposes a solution that elevates tactile sensing to the same level as vision, treating both as future observations to be predicted. Within the model's architecture, action generation can read estimates of future visual frames and future tactile contact during every denoising step. The most intuitive validation for this approach comes from high-contact-density tasks: in DexJoCo's assembly and tablet unlocking tasks, the robot must stabilize objects while performing fine-grained interactions, generating tactile feedback with both hands simultaneously. Publicly released rollout results show that ME-Dex 1.0 demonstrates more complete bimanual coordination than π0.5 and DECO on these two tasks.
02 One Representation for Five Hands
Tactile data and image data are not the same. Image formats are standardized, whereas tactile data formats depend on the sensors mounted on the robotic hand; changing the robotic hand yields a completely different set of data. The team's primary goal is to enable heterogeneous tactile data to be fed into a single model.
RoboTwin and ManiFeel use grippers, while DexJoCo and Xynova Flex2 employ dexterous hands. The physical robots are also equipped with Paxini PX6AX sensors housed within LeRobot SO-101 grippers. These sensing surfaces output a three-axis force field consisting of one normal axis and two tangential axes, with varying layouts, resolutions, and coverage areas. Preprocessing is required before the force field inputs: the team configures a layout mask for each sensing surface to mark the valid positions of the sensors; the force signals undergo smoothing and dead-zone processing first, then compress large values to complete normalization and clipping, with the normal component limited between 0 and 1, and the tangential components limited between -1 and 1.
The team uses the human hand as a template to build a hand-region model named Canonical Hand Model, treating various robotic hands as variants of the human hand form. The left and right hands have predefined corresponding regions for fingers and palms, with each phalanx and palm area of dexterous robots mapped to the corresponding positions on the template. The processing for RoboTwin grippers is more detailed: its two sensing surfaces are mapped to the thumb and index finger regions, each further divided into four ordered patches, ensuring that contact forces during gripper closure have fixed landing points.
Different hardware bodies can only cover partial areas of the template. The team uses area masks to distinguish between two scenarios: "this area has no sensor" and "there is a sensor but no contact at present." These two cases have entirely different implications in subsequent predictions: the former indicates that data cannot be obtained, while the latter signifies that a contact state has been detected as empty.

After completing spatial mapping and alignment, the force field must be compressed into a unified feature representation. The team trained a shared tactile autoencoder with parameters shared across all sensing surfaces, processing each frame independently to output a set of regional features carrying region markers. The encoder extracts local force features using residual convolutions, performs windowed attention within individual sensing surfaces, and applies masked pooling based on valid positions; for gripper sensing surfaces, pooling is conducted separately within four patches, while for dexterous hands, pooling is performed globally, followed by placing the results in fixed positions according to the template region order.
The decoder retrieves features from each sensing surface, uses local mesh positions as queries to perform cross-attention, and then restores surface features via residual networks and convolutions. The model is configured with two prediction heads: a force head predicts normal and tangential components, while a contact head predicts contact probability; the final restored force field is obtained by multiplying the two prediction results with an effective position mask.
Training this autoencoder uses four loss terms. The force loss supervises the prediction before contact gating at contact positions, and supervises the reconstruction after gating at background positions. The contact loss employs balanced binary cross-entropy, with labels derived from the magnitude of force values prior to dead-zone processing. The variation loss aligns force differences between adjacent frames, increasing weights for scenarios involving contact onset and release. The latent loss converges features in non-contact effective regions toward the result of zero input with the same layout. After joint training of these four losses is complete, the autoencoder is frozen to provide current tactile conditions and future tactile targets for the joint model.
RoboTwin and DexJoCo natively lack tactile sensing capabilities. The team built an automated data platform that replays existing action trajectories from both platforms, directly reads force sensor data from the simulation environment to generate tactile sequences aligned with images and actions, and simultaneously generates and executes new trajectories to expand the dataset. RoboTwin was supplemented with 50 tasks and 27,500 Clean and Random trajectories. DexJoCo added 11 tasks and 1,100 demonstration samples according to its official multi-task setup, including six single-arm tasks and five dual-arm tasks. ManiFeel features a three-dimensional tactile force field; for four insertion tasks, it uses 50 demonstration data points each to train policies separately.
03 Three Expert Paths for Joint Prediction
The main body of ME-Dex 1.0 is composed of three expert modules concatenated together, with video, tactile, and action each occupying one module; all adopt a denoising-and-reconstruction framework. The three modules retain independent processing pipelines while sharing attention at intermediate network layers. During each step of the denoising process, the action generation module can read predictions from the other two modules regarding future states.
The vision branch follows the Wan2.2-5B architecture, with 30 layers and a hidden dimension of 3072. The action branch has 30 layers and a hidden width of 1024; the tactile branch has 30 layers and a hidden width of 512. The vision and action branches use weight initialization pre-trained on large-scale robot datasets by Motus; the tactile branch uses random initialization. During training, the vision branch receives encoded first-frame images and future visual features with added noise; the action branch receives proprioceptive states and real action sequences with added noise; the tactile branch receives encoded current tactile frames and future tactile features with added noise. The ground truth for future visuals, actions, and tactile data serves as the supervision signal for each respective branch.
The tactile branch treats tactile data as a time series isomorphic to video. A future tactile sequence comprises T frames, with each frame describing the force distribution on both hands at that moment. The frozen tactile encoder processes these frames individually, outputting a feature sequence with spatial structure; the task of the tactile branch is then to reconstruct this set of features from noise, before passing them to the frozen decoder for reconstruction into a force field.
Information exchange among the three modalities is restricted to intermediate layers of the network. Drawing on the H-Bridge approach, the team projects features from all three streams into a shared attention dimension at these layers, performing unified attention calculations for visual, tactile, and action tokens before mapping outputs back to their respective original dimensions. Cross-modal information exchange concentrates in the middle layers; each modality retains independent processing logic during input and output stages, requiring fewer layers for cross-modal attention. This allows the action branch to read future predictions from vision and touch through just these few layers.

All three branches employ conditional flow matching for training. Each branch independently samples a noise level, adds the corresponding noise to clean future targets, and then instructs the model to predict the vector pointing from the noisy state to the target. The video and action branches use mean squared error; the tactile branch uses weighted squared error, assigning higher weights to frames with larger changes in tactile features. The three loss components are summed with weights and backpropagated together to update parameters.
During inference, future visual features, future tactile features, and actions all start from Gaussian noise and iteratively progress from noise level 1 to zero; the shared layer exchanges information at each iteration step. The action module serves as the control output, while the frozen visual and tactile decoders reconstruct the corresponding predicted frames and force fields.
04 Achievements Across Three Platforms
The evaluation results come from three simulation platforms and two sets of real-robot experiments. The evaluation metric is the execution success rate of the robot completing a single task in its entirety, with judgment criteria following each platform's official rules; multi-task experiments all use only a single training seed.
The performance gap is most pronounced on the DexJoCo platform. Under a multi-task training setup across 11 tasks, ME-Dex 1.0 achieved an average success rate of 65.3%, compared to 56.4% for DECO and 57.8% for DECO.p. Among five two-handed tasks, ME-Dex 1.0 reached an average success rate of 48.8%, while both DECO and DECO.p scored 39.6%. Looking at individual tasks, the model showed its greatest advantage in assembly and microwave oven tasks: baseline models achieved success rates between 0% and 6.0% for assembly, whereas ME-Dex 1.0 reached 28.0%; for the microwave task, performance improved from 64.0% with DECO to 88.0%. It also achieved the highest scores in the group for unlocking a tablet (48.0%) and the Tower of Hanoi (32.0%). In single-arm tasks, both grasping a water bucket and using pliers reached 92.0%. The gap was particularly significant in the pliers task, where DP-T scored only 6.0% and DECO.p reached 60.0%.
Insertion tasks represent another critical dimension of evaluation. In the four insertion tasks within ManiFeel, ME-Dex 1.0 achieved an average success rate of 70.0%, a 15.0 percentage-point improvement over the official baseline of 55.0%. The success rate for plugging in power cords rose from 58.0% to 88.0%, while the pin and USB insertion tasks each improved by 18.0 percentage points. Strategies for the four tasks were trained independently, with each task using 50 demonstration samples and 50 evaluations.

On the RoboTwin platform, ME-Dex 1.0 achieved success rates of 91.56% and 91.92% under Clean and Random configurations, respectively, when trained with a mixed dataset of Clean and Random data. These results exceed those of Fast-WAM by 2.34 and 2.70 percentage points, as well as those of Fast-WAM+Tac (equipped with tactile conditions) by 1.34 and 1.70 percentage points. The team conducted a layer-by-layer ablation study under this setup: connecting only the current tactile conditions to the visual-action model raised the Random task success rate from 86.90% to 89.54%; adding future tactile prediction further increased it to 90.60%; and finally, replacing the layer-wise joint attention with a mid-layer shared attention architecture reached 91.92%. This stepwise growth demonstrates that performance improvements stem from the combined effects of tactile supervision and cross-modal information exchange.
When trained on only 2,500 Clean demonstration samples, ME-Dex 1.0 achieved 89.6% in the Clean configuration and 68.1% in the Random configuration, with a mean of 78.9%, outperforming the strongest baseline OLA-Sem's 71.3% shown in the table.
In real-robot experiments, the team installed the Paxini PX6AX sensor into the LeRobot SO-101 gripper to complete the transfer of a power plug between two sockets; another experimental group used the XR5-10-FS robotic arm paired with the Xynova Flex2 dexterous hand to complete a yellow can stacking task. These two sets of real-robot experiments were conducted solely for qualitative validation, intended to confirm that the models could perform physical manipulation relying on visual and tactile observations. The study also separately tested the tactile representation module: the shared autoencoder achieved F1 scores of 96.56%, 90.05%, and 99.69% for contact detection on the RoboTwin, DexJoCo, and ManiFeel tactile datasets, respectively; the F1 scores for transition between contact occurrence and release states were 89.73%, 68.73%, and 66.96%, respectively.
05 Final Thoughts
The most core change in ME-Dex 1.0 is redefining the position of tactile sensing within the model pipeline. Tactile data is no longer just a conditional signal at the input stage; it becomes a prediction target on par with images. The action generation module can read estimated results of future contact states during every denoising step. This modification upgrades the logic of using tactile information from "reading current force" to "predicting multi-frame contact changes."
For dexterous manipulation, this approach provides reusable technical interfaces. By leveraging a hand template alongside a shared autoencoder, heterogeneous tactile information from grippers and dexterous hands can be unified into the same regional features; tactile data can also be automatically generated in simulation environments through trajectory replay. Contact-intensive two-handed tasks and insertion-type tasks show the earliest benefits, with ME-Dex 1.0 achieving a significant margin over baselines on these two task categories.
