Tsinghua’s Cao Ting on Physical Self-Evolution and Causal Learning for Embodied AI
On September 21, the ICCL Zi Jin Hua · Wu Li Zhi Neng Open Day was held at the Intelligent Industry Innovation Center of Tsinghua University Wuxi Research Institute. Zhang Ya Qin, an academician of Zhong Guo Gong Cheng Yuan and founding director of the Institute for Intelligent Industry at Tsinghua University, made a rare appearance to deliver remarks, introducing the concept of Physical Self-Improvement (PSI).

Academician Zhang Ya Qin posits that intelligence currently falls into three categories: mathematical, physical, and biological. He argues that the bottleneck in embodied AI lies in inconsistent data standards during the scaling phase and the lack of convergence in data scaling. To achieve higher-level physical intelligence, PSI will be the next key focus.

ICCL is the critical pathway to realizing this self-evolution and driving leaps in robotic capabilities. To bridge the gap between demonstrations and real-world application, Professor Cao Ting and her team at the Institute for Intelligent Industry at Tsinghua University have built an embodied self-evolution system comprising Model, Harness, and Infra.
ICCL (In-Context Causal Learning) serves as the core mechanism at the Model layer: robots extract causal relationships between actions and state changes from their own physical interactions, injecting these experiences into subsequent decision-making. Zetta Harness handles online supervision and immediate recovery during execution, as well as offline attribution, experience optimization, and skill accumulation after execution. Z-Infra provides full-cycle support for the model and harness, including data pipelines, environment rollouts, and efficient deployment inference. This physical self-evolution is not achieved by ICCL or any single module alone, but through the complete closed loop formed by the synergy of Zeva Model, Zetta Harness, and Z-Infra.

At the event, Professor Cao Ting and her team officially released the new generation self-evolving model Zeva-Ego, while outlining a clear roadmap for Z-Trans’s future vision. According to reports, Zeva-Ego achieves the extraction of physical causal priors from unannotated ego-centric data for the first time, marking the formal establishment of a dual-technology system consisting of offline 'causal priors' (learning from humans) and online 'in-context causal learning' (ICCL).
Why is causality so important? What is the relationship between self-evolution and 'causality'? Is there a more sophisticated way to combine them?
Through a major product launch, Z-Trans provided an answer. 42HOW Robotics interviewed Professor Cao Ting from the Institute for AI Industry Research (AIR) at Tsinghua University to reveal more details about ICCL and physical self-evolution.
Starting from the Bottleneck
'Causality' is an eternal proposition describing the complex changes in the physical world. From the chaos of reality, humans extract a path that describes these changes, using it as a foundation to understand and learn across time scales. Without a description of 'causality,' knowledge and experience struggle to become sustainable or transmissible. Similarly, for artificial intelligence to enter the physical world and develop higher-level general intelligence brains, designing a paradigm for causal learning is particularly crucial.
However, many large models we see today remain carriers of probability distributions. Huang Bi Wei, founder of Aether AI, summarized the evolution of AI into four paradigms: small correlation models, small causality models, large correlation models, and large causality models, with the first two still dominating the current landscape. Shifting from state trajectories and visual reconstruction to root-cause tracing, enabling AI to understand the causal laws of the world, may signify a deep shift in machine learning paradigms.
Professor Cao Ting believes that machine learning relying solely on probabilistic correlations is clearly insufficient for embodied models.
The current mainstream embodied learning paradigm relies on offline data collection to train models, which can only reproduce data seen during pre-training. Consequently, training via visual reconstruction is essentially a probability distribution of actions. Inferring human intent by learning these probabilistic associations does not enable machines to fully master human capabilities. Because the real world contains numerous physical conditions unseen during pre-training, pre-training alone cannot achieve generalized embodied manipulation. While embodied models may perform relatively well in demo scenarios with fixed contexts, backgrounds, lighting, and task objectives in the real world can change drastically in short periods; this learning method is far from capable of coping with millisecond-level changes. As a result, generalization can only infinitely approach the true distribution but never reach convergence.
Similar to human growth, stumbling steps are the first stage of standing in the real world, which is why causal interaction must hold in machine learning. She admits that the upper limit of embodied model evolution is currently hard to foresee. If evaluated by score, the capabilities of such models might be limited to a '60-point' level. Reaching 100 points remains a long and arduous journey.
The 'Cause and Effect' of Embodied Self-Evolution
Focusing on embodied self-evolution, Cao Ting’s team has successfully established a complete framework for physical agents. Prior to this, the industry was largely engaged in massive data collection and building data factories, hoping to expand intelligent capabilities by increasing data volume through pre-training. Despite this prevailing approach, Cao Ting’s team remained steadfast in its original vision.
"Thus, when the industry suddenly realized the critical importance of embodied self-evolution, we had already been releasing continuous results. People began to pay attention to it, which has both positive and negative aspects. On the less favorable side, various so-called forms of embodied self-evolution have emerged, but they may not align with our understanding. Therefore, today we clearly define what we mean by embodied self-evolution."
Embodied self-evolution refers to enabling robots to continuously interact with objects and their environment in the physical world, extracting correct and incorrect, good and bad causal outcomes from their own experiences, and gradually consolidating their capabilities; the more they are deployed in the real world, the better their self-evolution becomes.
If one were to compare this concept with the widely discussed RSI (Recursive Self-Improvement) in the industry, the defining characteristic of physical self-evolution lies in its emphasis on the embodiment itself. Taking autonomous driving as an example, physical AI has formed a deep learning paradigm by integrating traffic scenarios with vehicle platforms. The next step is to build underlying infrastructure that enhances machine labor power, allowing the embodied brain to rapidly establish data flywheels across different applications to realize core value, while also accelerating the replacement of human responses in scenarios such as disaster relief and high-mobility operations through an embodied intelligence platform featuring "one brain, multiple bodies."

Furthermore, PSI operates at a deeper level. The self-evolution of physical intelligence requires not only the automatic optimization and positive feedback loops of large models but also the formation of a top-down efficiency closed loop between reality and simulation, brain and body, and system and hardware. For instance, according to Anthropic’s latest report on the development of next-generation intelligent models, the proportion of work performed by machines increases over time, and the nature of improvements shifts from shallow tasks like optimizing prompts and modifying peripheral code to deeper optimizations involving attention operators, hardware architecture, and benchmark design.
This reveals that Academician Zhang Ya Qin’s explanation of PSI at the conference presents a rather vivid concept. However, reaching this vision requires both technical depth and horizontal expansion in physical AI.
Contextual causal learning represents another major innovation in the technical philosophy.
Earlier, Skild AI, Generalist, and Ma Yi Ling Bo released their respective embodied models in quick succession, revealing an industry consensus on In-Context Learning (ICL). However, ICCL (In-Context Causal Learning) differs from ICL. It posits that robots should learn in real-time from their physical interaction experiences during execution—specifically, the causal relationships between their actions and the resulting state changes—and apply these insights to subsequent actions. To this end, Cao Ting’s team published a related paper earlier this month, proposing the Zeva model trained via contextual causal learning.

Why is this unique?
Before causal learning, the default approach for training embodied models relied on associative learning. By inputting an image and outputting an action, the robot learns the correlation between video and movement. If tasked with opening a drawer, the robot receives an image of a drawer as input but merely executes a mechanically learned motion without understanding the impact of that action on the object itself. In contrast, ICCL informs the robot of explicit outcome changes: a pulling action might result in the drawer moving out 5 centimeters or 10 centimeters. This makes machine learning effects more tangible and grounded.
The model’s ability to absorb context is equally critical. Within the ICL paradigm, enhancing performance during each execution requires fully leveraging context. This includes not only human demonstrations but, more importantly, the robot’s own operational experience. Knowledge gained from the robot’s direct contact with physical objects and the environment is retained in memory. This connects Physical Self-Improvement (PSI) with contextual causal learning.

From a model architecture perspective, the application of contextual learning paradigms is at the forefront. The system achieves physical self-improvement through three stages: causal interaction extraction, dual-timescale causal memory, and context injection strategies.
One of the core distinctions lies in the encoder. To achieve Causal Interaction Extraction, the team designed a Causal Transition Encoder (CTE). During task execution, model parameters are frozen; only the extracted causal links between "action and state change" are added to the context. This approach places significant demands on the model's ability to learn autonomously within parameter constraints.
Professor Cao Ting recalled that this was a true process of going from 0 to 1.
Initially, the challenges were numerous. The team believed that while visual reconstruction encoders brought good results to pre-training, they could never truly achieve generalization. Although the concept of "causality" had been established as the cornerstone of the model's philosophy, the team had not yet fully figured out how to design this formal system. More importantly, there were no ready-made answers in the industry for this problem. After months of trying different routes, the design for the causal encoder was finally completed.

Zeva’s causal encoder is a complete innovation, falling outside traditional encoder paradigms. It differs significantly from conventional visual reconstruction encoders by integrating visual latent variables, action encoding, and observed outcomes into a Causal Interaction State. This state is then mapped to a Phase Token and a Causal Interaction Signal, forming a signal combination used for model training.
Subsequently, Dual-timescale Causal Memory organizes these signals to support adaptation during deployment. Brief Interaction Trace (BIT) captures dynamics within a single attempt, while Persistent Interaction Memory (PIM) accumulates useful evidence across multiple attempts within the same round. Once completed, In-Context Policy Injection retrieves interaction evidence matching the current phase to construct a Causal Prompt for the frozen base policy. During deployment, all model parameters remain frozen, with only the interaction memory updating online.

As interaction experience accumulated through repeated attempts, Zeva’s task success rate jumped significantly. Using a chemical laboratory as the scenario, Zeva achieved the best average success rate at every complexity level on the real-robot ChemLab-Evo benchmark. Compared to the strongest baseline at each level, it improved averages by 6.6, 5.0, and 5.0 percentage points at the atomic, short-sequence, and complex levels, respectively. Its average score of 57.32 exceeded the best baseline’s average by 12.59 points, surpassing models such as FastWAM and Cosmos3-Nano overall.

Zeva’s post-deployment continuous learning not only expands its capabilities but also demonstrates strong single-sample learning and cross-task generalization. Evaluations on ChemLab-Evo, using a fixed set of rounds, show that after four consecutive learning cycles, Zeva consistently improved task execution: "pick test tube" success rose from 65% to 100%, "place beaker" from 25% to 70%, and "pour liquid" from 30% to 80%. Across different tasks, Zeva exhibited robust cross-task generalization, further confirming that observed state changes enhance the discrimination of causal interaction memory.
Physical Agents: Growth of "Flesh and Blood"
To reach true physical intelligence, embodied self-evolutionary learning cannot remain confined to pre-training; it is crucial for robots to iterate through continuous self-interaction. Cao Ting’s team has first established a self-evolution framework comprising three parts: the model, infrastructure, and Harness.

As the training and inference infrastructure for Zetta, Z-Infra supports the entire self-evolution process, including real-time execution at the edge and response enhancement on the server side.
It must fully support the self-evolution cycle in two aspects. First, at the edge, it enables real-time inference where the robot is deployed. Second, on the cloud, it manages the specific effects of scaling. By unifying tools such as CPUs, GPUs, real robots, and physical simulation environments on the cloud side, Z-Infra decouples upper-layer models from underlying heterogeneous hardware. This allows upper-layer models to simply submit requests while various underlying hardware resources are managed separately.
This is no small feat. Physical agents differ significantly from digital ones. Digital agents for scenarios like gaming or information exchange can be implemented with code, skill libraries, and knowledge documents alone, whereas physical agents constitute a systematic engineering project spanning models, algorithms, agents, and hardware. For physical agents, managing complex computing resources remains the greatest challenge.

When a task starts, physical agents typically invoke interaction trajectories (rollouts) in large-scale parallel. This approach requires an efficient combination of diverse computing resources. For CPU-centric simulations, environments like MuJoCo and Robosuite maintain simulation states per session, consuming the host machine's CPU and memory. Meanwhile, other functions are distributed across GPUs, which have varying requirements for state layout, rendering, and video memory. Since policy models use GPU accelerators for autoregressive or diffusion-based action generation and require video memory to store computational features, these workloads impose different demands on workers and scheduling. Consequently, no single resource pool or scheduling strategy can accommodate all these workload types.
This places dynamic coordination requirements on the team. Unlike training pipelines with predictable data-parallel patterns, agents dynamically select which tools to invoke based on observations: one action may require only a single policy call, while the next step might chain perception, planning, and multiple primitive operations, making trajectory invocation difficult to predict. As agents explore, reflect, and retry with modified policies, sessions are created, paused, and destroyed at irregular intervals. This leads to bursty GPU inference demands, with significant variance in batch sizes and arrival intervals, rendering static resource allocation and request scheduling inefficient.
"At that time, our algorithm team, Infra team, and Agent team held several meetings, and there were debates on this issue. It was essentially because everyone was considering points from different levels. Although communication between systems and algorithms was frequent, due to their respective roles, everyone focused on maximizing efficiency within their own domains, and their ways of thinking were sometimes quite different."
Fortunately, this is Cao Ting’s team’s area of expertise.
"We already had extensive technical accumulation in this area, such as low-bit quantized model deployment, unified KV Cache management, and device scheduling during task concurrency. Previously at Microsoft, our group was responsible for enabling Microsoft large models to run at high speeds on edge devices like mobile phones and PCs under limited hardware resources. In system optimization, we integrated all these technologies, so ultimately they were all applied effectively."

Finally, the team chose to decouple the upper-layer algorithms from the lower-layer hardware. The algorithm team focuses solely on rollout completion rates without worrying about GPU compute resource distribution; the systems team handles and dispatches robot requests overall, coordinating compute resources. Workers responsible for environments mainly manage simulation environments, sessions, and physical states, while another set of Rollout Workers handles VLA/WAM model inference and asynchronous scheduling, supported by various underlying hardware. With each party performing its duties, Z-Infra’s architecture became clear.
Robots, and How to Reflect on Closed-Loop Control
With foundational infrastructure and "brains" in place, the engineering "scaffolding" that connects them cannot be overlooked. Harness is responsible for monitoring whether models can execute tasks correctly and continuously. From Claude Plays Robots to GPT-6 Astra, which can also command embodied carriers to perform operations, the connection between large models and Harness engineering is growing tighter and more closely aligned with application scenarios.
However, entering the physical world for embodied interaction is not easy. In recent research on physical intelligence, one path involves training end-to-end policy models for general robot control. While promising, real-world deployment has exposed gaps such as expensive data requirements, drift in physical distributions, and the cascading amplification of minor execution errors over long-horizon tasks. These issues limit the progress of end-to-end models from demonstration to real-world implementation.
The second path explores embodied agent capabilities, filling these gaps by surrounding base policies with LLMs, code, tools, memory, planning, verification, and recovery. Early systems like PaLM-E and Code as Policies demonstrated the potential of language models in embodied reasoning and planning, while RoboCat showed that agents could expand their operational capabilities by collecting data from their own attempts across tasks and embodiments. Recent demonstrations from systems including Claude Plays Robotics, CaP-X, HarnessVLA, and Guava further indicate that coding agents and execution frameworks can make frozen or pre-trained robot policies more reliable, without waiting for end-to-end strategies to be fully resolved. This reveals a shift in the current technological frontier: moving from relying solely on single end-to-end policies toward coordinating scaffolding systems that integrate policies with planning, memory, tools, verification, and recovery.
The design complexity of embodied agents increases further. Because they operate in continuous environments, Harness must simultaneously evaluate both the robot's state and the world's state, deciding when to invoke policies, tools, evaluators, or recovery modules, and preventing early errors from evolving into unrecoverable physical failures. Additionally, given that current VLAs are significantly weaker than frontier LLMs in terms of capability and reliability within their native digital domains, embodied agents inevitably require more external tools to support robust task execution. Consequently, many embodied agents remain "turn-based": they may achieve recovery within a single trial but rarely convert execution trajectories into governed, long-term improvements.
Limitations also stem from hardware. Due to the massive parameter size of large models, it is impossible to efficiently fit them entirely onto edge devices, resulting in significant inference latency at the embodied level. Since real-time invocation of cloud-based large models for millisecond-level responses is not feasible, only limited post-hoc attribution is possible: sending the complete trajectory of task execution to a large model for analysis and review.

Where is the breakthrough?
Zetta takes a different approach to physical intelligence. While keeping its foundational policy model frozen, it evolves an online, code-based runtime critic and recovery skills, bridging the gap between static open-loop agent scaffolding and the high-frequency governance required for physical execution. Through three closed loops—action-level error correction, batch experience optimization, and gated validation updates—the agent achieves online error correction and skill accumulation during long-horizon tasks, enabling robots to continuously correct deviations and accumulate experience in complex scenarios.
The Code-based Experience framework, first proposed by Cao Ting’s team, allows the Zetta agent to perform action-level supervision purely through code. Data collected after execution is sent back to the server for further improvement of both the Harness and the model itself. When the model makes an error, the Zetta agent intervenes proactively, analyzes the cause of the failure in the Harness, and generates very lightweight Critic code suitable for local execution. This code is stored alongside context as memory; if the agent gets stuck, this triggers supervisory functions for rapid fault tolerance and task completion.

Taking the example of picking up a water cup: initially, the Z0 robot failed to grasp it successfully, crushing the cup due to excessive force. The model then entered a failure feedback ReAct Loop, readjusting the grasp position, clamping force, and motion trajectory. After multiple rounds of causal reflection and action verification, the Z0 robot could complete the task without damaging the object, writing this experience into the tool memory library via Scene Memory.

Deploying the Zetta physical agent on LIBERO-Pro resulted in significant improvements in success rates. Compared to other state-of-the-art (SOTA) models, Zetta achieved average success rates of 92.5% and 89.0% across different task categories on LIBERO-Pro, with scores of 63.0% and 40.0% under the corresponding LIBERO-10 settings. Zetta demonstrated superior performance compared to the π 0.5 model.
For Zetta, learning to 'reflect' is a defining characteristic. During task execution, rather than experiencing action drift, Zetta’s precision improved with each learning round, leading to an 'Aha Moment.' The agent reflects on potential errors in every step it takes, allowing the robot to continuously adjust and understand how tasks should be executed in terms of motion and precision. Breaking away from post-hoc attribution is key to Zetta’s formation.
"Previously, many others tried having robots generate successful or failed trajectories and feed them to large language models for self-improvement through reflection, but the results were unsatisfactory."
In fact, this post-hoc attribution approach is akin to carving a mark on a moving boat for machine learning. Because reconstructing failure scenarios and behaviors comprehensively through retrospective analysis is generally difficult, large models often rely on guesswork to learn, significantly reducing efficiency. In contrast, the model deployed on the Z0 robot emphasizes online self-learning; the agent reflects on potential errors in each step, allowing the robot to continuously adjust and understand how tasks should be executed with the right actions and precision.
Cao Ting’s team employs a fundamentally different error attribution process. For the model, it can update gradients to approach the optimal solution within the optimization space. The supervised signal learned by Zetta is a function or skill that reconstructs the entire process via simulation, ensuring that success rates continue to rise.
"After obtaining the execution trajectory, we first group it. The purpose of grouping is to categorize the causes of errors, preventing minor differences from causing misunderstandings during the large model's attribution analysis. We then sort them to identify common, severe issues that could lead to subsequent failures. Next, we propose corrective methods. After these are proposed, we return to evaluate whether the methods are effective."
After multiple rounds of error correction and iteration, this method is recorded in the supervision library as permanent memory retained within the agent. Meanwhile, the routine updates to the skill library adhere to strict standards, ensuring that random performance in a single round does not contaminate the skill library.
"It is like when humans make a mistake: we analyze where we went wrong today and try things out—'Oh, maybe there is a bug here, another there.' Only after learning and deducing for a while do we have an epiphany: 'Ah, the problem was actually here; this is the core issue.' When the machine identifies the root cause, success rates soar. This is similar to why large language models were said to experience an 'Aha Moment'—the sudden realization that previous errors were caused by this specific reason."
Behind the New Zeva-Ego Model

The Zeva-Ego model, officially released at this conference, stands out for its first use of unlabeled ego-centric data to extract physical experience. Cao Ting’s team utilized 10,000 hours of unlabeled first-person-view data to enhance generalization, formally merging offline causal experience and online context into a self-evolving dual-track system.

Zeva-Ego uses π0.5 as its vision-language-action (VLA) foundation model. The VLM handles scene understanding, task comprehension, and prediction of the next sub-task, while the action expert generates specific continuous actions for the robot. Cao Ting’s team achieved a breakthrough by efficiently extracting causal relationships from Ego data without requiring labeled datasets.
"We added an Action-Centric Encoder to the front of the new model. It learns causality not only from the robot's own actions but also from first-person Ego operational views. This approach fully leverages the data efficiency of heterogeneous data sources."
The biggest difference between Zeva-Ego and the previous generation Zeva model lies in the richness of training data. In practice, human-collected Ego data is used for mid-training alongside other labeled and analytical data, building on existing pre-trained language models or VLA models to enhance causal reasoning through mid-training. Subsequently, post-training is performed using task-specific data. During inference, Zeva-Ego is restricted to learning contextual causality solely from the robot's own real-world physical contact.
This method has resulted in remarkably high data absorption efficiency for Zeva. Approximately 4–5 hours of heterogeneous data equates to just 1 hour of real-robot data, marking the first validation of Ego data learning efficiency. Consequently, the success rate in self-owned real-robot experiments rose from 0% to 60%. Furthermore, because Ego data can be used in mid-to-late stage training without annotation, the cost of collecting real-robot data dropped significantly.
Professor Cao Ting revealed that after the feasibility of contextual causal learning was validated with the Zeva model, development of Zeva-Ego was quickly prioritized internally, with both projects running in parallel at one point. Zeva-Ego further enhanced Zeva’s ability to reason about physical causality. By incorporating broader and more accessible data types into training, intelligent performance improved substantially. As Ego data volume grew from 0 to 10,000 hours, RoboTwin’s success rate increased from 63.8% to 75.3%. Compared to the π0.5 model, the real-robot success rate after Ego post-training showed a decisive lead. By combining human and machine data, Zeva-Ego extracted the most fundamental causes behind failures and successes. This indicates that as data grows, the model refines experience, leading to effective emergent intelligence.
"If we use only 10 real-robot data points, the success rate is 0%. However, adding just 40 Ego data points boosts the success rate immediately to 60% or 70%. The learning efficiency is extremely high."
Beyond efficiency, Zeva-Ego’s approach validates a route that also significantly lowers data costs. Starting with first-person videos from humans, a single person can continuously collect data using a camera during daily work. This is far less expensive than methods requiring the robot body and teleoperation equipment.
"We found that human operators are extremely fast; however, teleoperation requires purchasing robot bodies and various control devices, which is very costly. By significantly reducing the amount of robot data used and increasing the proportion of human operation data, we discovered that the results were actually better."
Professor Cao Ting stated that the team has prepared a long-term roadmap for Zeva, and all related outcomes have achieved ideal results as expected.
"Once Zeva was successfully implemented, Zeva-Ego accelerated significantly because our self-evolution paradigm was fully established. Subsequent challenges were minor, such as improving the encoder's effectiveness when processing Ego data from real human operators."
To enhance robustness against interference, the team adopted a two-stage learning approach during Zeva-Ego training. Stage 1: The team explicitly separated action components from environmental interference using labeled action data. When the camera was disturbed by the operator’s own movements (e.g., nodding or shaking the head), this interference was isolated into an environmental channel, explicitly distinguishing "intentional behavior" from "unconscious camera movement." In Stage 2, the team processed unlabeled Ego data through the encoder trained in Stage 1 to separate irrelevant human movements from genuine hand and arm actions, thereby eliminating interference.
"Removing interference from human Ego data while effectively incorporating it was the biggest challenge we faced while advancing Zeva-Ego."

In practical evaluations, Zeva-Ego consistently outperformed other models, including π0.5 and Fast-WAM, across multiple task metrics in RoboTwin 2.0 (Hard).

Expanding Ego data enhances model performance and strengthens generalization. In a chemical laboratory scenario, for example, the current Zeva-Ego can already complete long-horizon tasks such as adding test tube solutions and performing pH titration tests. This accelerates the large-scale deployment of Z-Trans’s embodied AI models.
Building Infrastructure for the Machine Labor Era
Z-Trans has a clear technical vision. As a frontier technology startup incubated by Tsinghua University’s Institute for AI Industry Research (AIR) for the era of physical intelligence self-evolution, Z-Trans is dedicated to building foundational technologies for "embodied self-evolution," enabling robots to become stronger with use in the real physical world and driving the industrial adoption of embodied AI from "capable of demos" to "capable of reliably completing tasks continuously."
With the release of the Zeva-Ego model, Z-Trans continues to address challenges in cross-embodiment generalization, long-horizon task execution, experience feedback loops, and real-world closed-loop operations, further advancing the self-evolution of physical AI. Within the ICCL paradigm, the Zeva model still has significant room for evolution.

Cao Ting believes that the main difficulty in long-horizon tasks lies not only in duration but also in the need to gradually improve the success rate of atomic operation combinations through enhanced supervisory signals. To this end, Z-Trans plans to release optimization models targeting specific deployment limitations, along with additional related technical achievements.
There are also emerging voices in the industry recently. As data volume increases, differences between embodiments have become less pronounced. This is why we place Ego data in the middle stage of training. Previously, the industry discussed an embodied data pyramid: first training on human operation Ego videos, then on robot operation data, and finally on teleoperated robot data specific to the current scenario. Data effectiveness was considered higher toward the top of the pyramid. However, it has recently been realized that when robot-specific data becomes sufficiently abundant, the model learns more generalizable manipulation skills. At that point, the specific embodiment itself becomes less critical.

Regarding Z-Trans’s primary objective, Professor Cao Ting stated that the company has completed its technical chain verification and is currently following a clear long-term roadmap. Using the laboratory as an example, she explained that Z-Trans’s self-evolution process must first achieve end-to-end success in a single point within one scenario, forming a complete closed-loop solution in high-value scenarios to establish benchmark deployments. Subsequently, the model’s algorithmic capabilities across applications must ensure generalization in different contexts. Finally, by increasing the number of deployed robots, embodied AI can operate long-term and stably across the entire infrastructure.
Beyond technology, Z-Trans aims to build robust commercial barriers. From models and agents to infrastructure, the robot’s ‘brain,’ skill libraries, and cross-scenario generalization capabilities will form a business model centered on robotic labor. Relying on the PSI (Physical Intelligence Self-Evolution) system, and leveraging ecosystem partners from Tsinghua University AIR and the newly established Apollo community, Z-Trans plans to accumulate significant data assets through real-robot deployment and scenario expansion. This strategy positions the company to coordinate and complement with model providers and hardware manufacturers, securing a favorable ecological niche in the industry.
Through a recent collaboration with Tsinghua University’s Department of Chemistry, Z-Trans is bringing Auto Research into real-world scientific research scenarios. The system enables robots to perform tasks such as pipetting, weighing, and extraction, replacing student labor and freeing researchers from tedious work to automate the scientific process.
‘This also reflects our higher vision: to accelerate the development of scientific experiments and achieve greater scientific breakthroughs.’
For the infrastructure of the robotic labor era, Z-Trans’s causal logic and future prospects are worth anticipating.
