Dialogue with Ri Mian Kai Wu’s Zhong Yuan Xin: Industrial Settings Are Better for Validating Physical RSI

Recently, RSI has re-emerged as a high-frequency term in the AI industry. In embodied AI, it is being directed toward a more concrete question: Can machines improve themselves in the real world?

RSI, or Recursive Self-Improvement, refers to a system’s ability to participate in improving its own code, tools, or training processes, and to use these improvements to continuously enhance the capabilities or efficiency of subsequent iterations. As large models begin writing code, running experiments, and assisting in training next-generation models, this once largely theoretical problem is becoming increasingly realistic.

In the physical world, however, this closed loop is far more complex. Robots face continuously changing environments that cannot be fully reset; every action carries real costs and leaves real consequences. Physical RSI aims to solve how robots can do more than just execute tasks—how they can continuously gain experience from real-world interactions, update their capabilities, and bring those capabilities back into reality.

A robotics company founded about six months ago is attempting to bring this approach into factories right away.

Ri Mian Kai Wu was established in March 2026. In the following months, the company completed two consecutive seed funding rounds, totaling hundreds of millions of RMB. Investors include Ding Feng Ke Chuang, Yuan Tu Wei Lai, Bai Du Feng Tou, Wo Yan Zi Ben, Wu Yue Feng Ke Chuang, and Wan Lin Guo Ji. According to information disclosed at the time, these funds will be primarily used to build a self-developed world model, LaMPA, reinforcement learning systems, data closed loops, and product delivery capabilities.

Around the time of the funding disclosure, Ri Mian Kai Wu began releasing a series of technical updates, clarifying its technology stack.

On July 1, the company released frameworks for the World Reward Model (WRM) and World Policy Model (WPM): the World Policy Model generates initial actions, while the World Reward Model determines whether a task is successful, ongoing, or failed, allowing the robot to continue training based on real feedback. On July 8, the team further explained the underlying LaMPA world model and its triple representations: Environment, Ego, and Experience. On July 14, the company announced a VLA inference optimization scheme that reduced latency from 130 milliseconds to 32 milliseconds.

At this stage, what RimCrown describes remains a relatively standard physical foundation model system: the world model handles understanding and prediction, the policy model drives action, the reward model evaluates outcomes, and real-world reinforcement learning further refines the policy. The entire system is deployed on robots via inference infrastructure.

After August, the focus shifted. The question was no longer just 'how to train a better model,' but whether robots could continue learning after deployment, improving faster with each iteration.

This gave product delivery another layer of meaning. Robots entering real-world sites are not just for completing projects; they also serve as entry points for continuous interaction, failure, and feedback data.

On September 15, RimCrown officially released Physical RSI: Learning to Improve, in the Physical World, integrating prior technical work on world models, evaluation, post-training, and infrastructure into a single framework. It introduced two custom grading systems: AL0–AL3 and EL0–EL3.

In this framework, learning is split into two nested closed loops: the Skill Learning Loop converts demonstrations and trial-and-error into specific skills, while the Deployment Loop feeds issues exposed during real-world deployment back into the evaluation and training environments. RimPilot, RimCritic, and RimGround handle policy execution, outcome evaluation, and training environment construction, respectively. Ri Mian aims to ultimately update not only the robot’s action policies but also the evaluators, training environments, and even the learning process itself.

The next day, this newly named technical narrative entered an ICT testing production line at Yuan Tu Wei Lai in Dongguan.

On site, a robot had just finished operating on one type of circuit board, and staff immediately swapped it for another. The team collected a small amount of new data, completed training and deployment, and within about 15 minutes, the robot began processing the new materials.

Swapping a board poses almost no new problem for humans. For traditional automation equipment relying on fixed trajectories, however, a new board often means reprogramming, alignment, and debugging, which can take an engineer several hours. Ri Mian calls this integrated hardware-software system a 'super workstation' and summarizes its rapid changeover capability as 'learn once, apply everywhere.' According to company disclosures, this system has already completed a proof of concept (POC) on Yuan Tu Wei Lai’s ICT testing line, with a success rate exceeding 99.5%.

But the ability to learn something quickly and the ability to do it independently remain two different things.

We observed on-site that staff members still stood by workstations, participating in data collection, training, and task switching. Zhong Yuan Xin, co-founder and head of the technology platform at Ri Mian Kai Wu, did not shy away from this reality: 'It is not yet fully unmanned, so we are still a distance away from the true endpoint of RSI.'

Zhong Yuan Xin, born in 1997, majored in Vehicle Engineering at Tsinghua University for his undergraduate degree before earning a Ph.D. from the University of Michigan. During his doctoral studies, he interned with both Waymo and the Apple Vision Pro team. After returning to China, he joined Huawei’s Noah’s Ark Lab, where he experienced successive generations of intelligent driving technology roadmaps including end-to-end models, VLA (Vision-Language-Action) models, and world models, ultimately leading the development of the Huawei ADS 5.0 intelligent driving world model. He left Huawei at the end of 2025 and co-founded Ri Mian Kai Wu in 2026.

According to Ri Mian Kai Wu’s own classification, the capabilities demonstrated at Yuantu still belong to AL1: humans set the tasks, rewards, and training environments, while the system is responsible for collecting interactions within them, updating policies, and adapting to new tasks. The further step, AL2, requires the policy, evaluator, and training environment itself to begin updating collaboratively, proving through continuous new tasks that the time required for robots to reach goals is decreasing with similar data and training budgets.

In other words, what Physical RSI truly cares about is not whether a robot can relearn how to handle a circuit board in 15 minutes, but whether the cost of learning the next board decreases after it has learned the first, second, and third boards.

This corresponds to two paths being explored by robot foundation models: one seeks to give robots strong priors and generalization abilities upon entering a new environment for the first time through large-scale pre-training; the other allows robots to use real-world feedback to learn quickly after entering a specific environment. Figure’s Helix 2.5 emphasizes the former, while Physical Intelligence has demonstrated the latter in fine manipulation—using real operation results for online reinforcement learning.

Corona chose a more realistic entry point focused on actual delivery. It did not pin all hopes on zero-shot generalization, but accepted the reality that, at least today, robots often still need to relearn when entering a new workstation. What it truly aims to change is whether this learning process leaves behind data, evaluation methods, and training infrastructure to serve subsequent deployments.

This also explains the importance of Yuantu. Compared to autonomous driving, which possesses large-scale real-road data, the robotics industry lacks a sufficiently large hardware network. Yuantu opened up real workstations, production processes, and internal task scenarios to Ri Mian Kai Wu, allowing real tasks and failures to continuously enter the training system.

This is also the distinction Ri Mian Kai Wu attempts to draw from traditional automation integrators: in traditional projects, acceptance often signals that delivery is nearing completion; in its envisioned closed loop, delivery is instead where learning begins. New materials and tooling expose failures, evaluation models judge the results, failure data returns to the training system, and new strategies are formed.

If this cycle holds, a robotics company’s assets are no longer just measured by how many workstations it has delivered, but also by the collective learning capacity accumulated across those stations.

Zhong Yuan Xin stated this more directly in the interview: 'Algorithms are not a moat. What truly cannot be taken away is data and integrated infrastructure.'

Below is an edited conversation between 42HOW Robotics and Zhong Yuan Xin, co-founder and head of the technical platform at Ri Mian Kai Wu.

Large Models Handle Thinking; Embodied Companies Must First Perfect the 'Cerebellum'

42HOW Robotics: Your last interview was in July, when you said embodied AI companies must train their own large models. With the recent release of GPT-6 Astra, attention has shifted to significant changes in upper-layer foundation model capabilities. Has your judgment changed?

Zhong Yuan Xin: It has become clearer. We originally envisioned a layered architecture, we simply did not elaborate on it in July.

Tasks leaning toward general knowledge and common-sense understanding should be delegated to GPT-6, or at least vision-language models with 30B or 100B parameters. For example, understanding what to grasp in a home environment requires not just object and material recognition, but also broad knowledge and extensive semantic data. Embodied AI companies have no need to retrain a model or reinvent the wheel to access these capabilities.

The top-level model should come from foundation model companies. Unless you are Alibaba and can use your own model, companies like ours will integrate other foundation models, performing at most external harnessing. Our focus is on embodied intuition System 1: generating timely feedback and adjustments when the robot contacts or interacts with objects. Below that lies System 0, which leans toward motion control.

System 0 and System 1 are the directions embodied AI companies should prioritize at this stage.

42HOW Robotics: Do you think the progress of large models meets your expectations?

Zhong Yuan Xin: It’s faster than I expected. I originally thought it would take until next year to reach this level of capability, but it arrived this year. If a foundation model company were to connect a robot body to perform manipulation tasks, I wouldn’t be surprised.

However, they likely won’t deploy such massive models, so real-time, high-frequency closed-loop control may not be their goal. In industrial scenarios, during the final moments of fine manipulation, large models often lack clarity. The encoder in large models is more oriented toward semantic understanding and hasn’t learned fine-grained alignment as well; this is precisely where we need to contribute.

I previously considered that after optimizing the lower layers, the upper-layer task decomposition could also be slightly fine-tuned. Now I feel it’s unnecessary. Using the large model directly for the upper layer, coupled with an external harness, suffices.

42HOW Robotics: If you only focus on System 0 and System 1, can you still be called a foundation model company?

Zhong Yuan Xin: Of course. Foundation models can exist for both System 1 and System 2.

The LaMPA model we discussed in July, or the physics-native world model, remains in System 1. It aims to build physical intuition: upon seeing an object, the system knows its approximate shape and how to grasp it; if a grasp fails, it can immediately provide feedback and adjust. It is called the "cerebellum" because reactions must be fast. However, this still requires pre-training on large amounts of data to form general capabilities.

42HOW Robotics: How large does the final System 1 model need to be, and how much compute will it require? Many assume edge-side models do not need to be very large.

Zhong Yuan Xin: Because they need to be deployed at the edge, our current assessment is that the model will not exceed 10 billion parameters (10B); anything larger would be difficult to deploy.

For a 10B-level model, even with 300,000 hours of training data, the training resources would likely range from one thousand to several thousand GPUs, depending on the training cycle. Our target is also at this scale.

42HOW Robotics: So, embodied AI companies do not need to train models as large as cloud-based foundation models.

Zhong Yuan Xin: Correct. The final form may involve providing the upper-level brain with a set of Skills. These Skills are complex and rely heavily on understanding physical interactions.

We explored similar approaches when working on autonomous driving. Vehicles need to see both wide-angle and telephoto views simultaneously: the wide angle handles understanding surrounding vehicles and their relationships, while the telephoto view clarifies traffic signs hundreds of meters away. A single model struggles to handle both well, so a large model can call an "amplifier": once it detects a potential traffic sign in the distance, it crops and zooms in on that part of the image.

Embodied AI will see similar division of labor. Grasping objects or operating microwave buttons can become Skills requiring fine-grained interaction; long-horizon tasks should not be delegated to the underlying small model.

42HOW Robotics: Will this change the valuation logic for embodied AI "brain" companies? Investors typically assign high valuations to model companies, expecting them to eventually dominate the market.

Zhong Yuan Xin: The future market will likely be split into at least two parts: one for the higher-level System 2 brain, and another for the System 1 cerebellum.

Can the cerebellum achieve winner-takes-all status? I think it’s possible. If our skills are stronger than competitors’ and we iterate faster, we have a chance to do so.

42HOW Robotics: Both OpenAI and domestic foundation model teams are purchasing ego-robot data. With similar data sources and model architectures, where will the differentiation among embodied AI companies lie?

Zhong Yuan Xin: The key lies in what exactly you are training: joint angles or virtual spatial representations; single-frame images or continuous video streams; or vectorized representations? There are no definitive answers yet.

Embodied AI companies can still explore representations more closely aligned with their specific hardware. We have experimented with data formats such as Ego and UMI, but it now appears that no single solution covers all needs. Data strategies must be tied to the final delivery method: Are you aiming for zero-shot deployment directly onto customer hardware, or do you allow minute- or hour-level fine-tuning, as we do, relying primarily on on-site calibration? Different delivery models require different pre-training data compositions.

If you pursue cross-hardware zero-shot capability, pre-training may require more data tightly coupled with specific hardware to minimize the hardware gap. If few-shot adaptation is acceptable, pre-training may not need as much hardware-specific data, and ego data can play a significant role.

42HOW Robotics: You refer to your representation system as E/E/E. Have you determined precisely what data needs to be trained at the operation layer yet?

Zhong Yuan Xin: We cannot disclose the specific technical details yet, nor have we published any papers. We have identified certain nuances that the industry is currently overlooking, and our strategy is to develop these solutions first before making them public.

The broader direction remains fine manipulation: determining exactly which modalities are required, what input frequencies are necessary, and whether data should be captured in Ego, UMI, or other formats. We are still analyzing this, but the final approach will likely fall somewhere between Ego and UMI.

Recently, after discussing potential collaborations with researchers in Singapore, we have become increasingly clear that fine manipulation tasks involving servers require robust force feedback. While some household tasks may not demand this level of precision, force feedback is almost certainly a necessity in industrial settings. Consequently, we are also designing data collection devices equipped with force feedback capabilities.

42HOW Robotics: You mentioned earlier that Yuantu aims to scale from hundreds of units to tens of thousands. Do you intend to pursue this scale-up exclusively through server manufacturing?

Zhong Yuan Xin: That is a misunderstanding. The goal of scaling to tens of thousands of units does not mean we are partnering solely with Yuantu.

We established a joint laboratory with Yuantu to co-develop solutions, which we then deliver to a broader customer base. Yuantu provides numerous internal scenarios for us to validate our methodologies. Once validated, what we package and deploy should not be limited to an AI-server-specific solution, but rather a system capable of generalizing across different scenarios. Our future customers will extend well beyond Yuantu.

Failure Is Easier to Transfer Across Embodiments Than Success

42HOW Robotics: You previously stated that data collection should be driven by delivery requirements. However, if foundation models are to establish industry consensus, they theoretically require universal data. Starting from delivery needs—does this represent only the first phase?

Zhong Yuan Xin: I will first distinguish between To B and To C.

To C data naturally possesses greater diversity. However, given current collection methods, the generalization capability of To B data is often insufficient. Whether in highly standardized industrial scenarios or hotel cleaning, there are uniform standard operating procedures (SOPs). When humans perform the same task stably according to standards, action diversity naturally disappears. If only correct behaviors are collected, continuing to collect data by switching to another industrial scenario may yield limited gains.

We now prefer to collect failure data. Successful actions tend to be similar, whereas failures vary in their specific 'misfortunes,' and a single scenario can present many types of failures.

Furthermore, we believe that achieving generalization at the evaluation level is easier than at the policy level. By observing an operation video, one can judge whether an action is correct based on whether an object drops or exhibits abnormal deformation, regardless of whether the executor is a human hand, a dexterous robotic hand, or a gripper. However, enabling the same policy to control both a dexterous hand and a gripper is significantly more difficult.

Therefore, our paradigm is: The policy does not necessarily need to achieve Zero-shot performance initially; Few-shot is acceptable. However, the model responsible for evaluation should strive for Zero-shot capabilities to cover more tasks. As long as it can stably judge the quality of actions, even an immature incoming policy can be trained quickly.

42HOW Robotics: In other words, you are not first insisting on unifying underlying action representations, but rather adding an evaluation system that can span across tasks and embodiments?

Zhong Yuan Xin: Correct. Once an evaluation system is in place, the importance of whether the underlying layer trains joint angles or other features diminishes. These choices affect the starting point and qualities of the 'athlete,' but the 'coach' responsible for evaluation can possess stronger generalization capabilities first.

42HOW Robotics: Can this approach ultimately unify To B and To C?

Zhong Yuan Xin: I haven't decided yet. It can certainly be used for To C applications, but whether it is the most suitable paradigm for homes or capable of generalizing across different households remains uncertain.

42HOW Robotics: Where does the failure data come from? Asking humans to deliberately perform failures makes it difficult to cover the errors robots actually make.

Zhong Yuan Xin: If we focus on delivery, obtaining failure data is not difficult. Once a scenario achieves a high success rate and moves to the next, new failures naturally flow back. Even within the same scenario and task, deploying in a different location will generate many failures.

For companies that only train foundation models, failure data is hard to obtain; but for teams close to actual scenarios, this is precisely an advantage.

42HOW Robotics: You have been working on world models for a long time. Why did you later develop the World Reward Model and World Policy Model?

Zhong Yuan Xin: We initially worked on world models to avoid heavy reliance on annotations. As long as there are future frames, we can train a world model. Looking at it now, we may still need to add some simple annotations, but it at least reduces annotation costs rather than requiring extensive manual labeling every time a new batch of data arrives.

Continuing our work, we found that the Reward Model is easier to generalize and migrate across different robot bodies and models. The World Model can reduce annotation costs, but directly training a Policy still requires consistent operational paradigms and end-effector representations. The World Reward Model is more pure; it is only responsible for understanding and evaluation, outputting a score or a brief description, without needing to concern itself with what shape the end-effector takes or its joint positions.

Of course, how much help the world model provides to the Reward is something we are still exploring.

42HOW Robotics: You have worked on world models for autonomous driving. What is the fundamental difference between a world model for autonomous driving and one for robotics?

Zhong Yuan Xin: On a smaller scale, the backbones used are quite similar; many teams rely on Wan at the foundational level. However, the core problems each field faces are different.

Autonomous driving must first solve multi-view consistency: how to ensure multiple cameras generate consistent futures. Wan was originally trained primarily on single views, so adding multi-view capability is the first major challenge. Embodied AI, by contrast, often needs to address how to output actions even earlier in the pipeline.

I believe embodied world models are more difficult. Vehicle trajectories are relatively constrained; in straight-driving videos, a model might learn seemingly good results simply by gradually enlarging distant pixels—a kind of shortcut. In robotics, external camera perspectives change significantly, and during manipulation, the robot constantly encounters areas it has never seen before. Directly applying video models creates problems.

It may require selective modeling. For areas outside the current view, some inference is necessary, but details like the color or material of an irrelevant wall do not matter. The model should allocate more capacity to manipulating objects and their immediate surroundings rather than generating the entire scene uniformly.

42HOW Robotics: Is a world model a solution closer to the ultimate goal, or just a usable tool borrowed from other fields for the present stage?

Zhong Yuan Xin: Autonomous driving is very pragmatic. When a new method emerges in another field, people try it; if it solves the problem, they use it; if not, they drop it. Meta-learning, causal inference, and neural-symbolic networks have all been tested. End-to-end systems, VLAs, and world models were adopted because they proved useful after testing.

The embodied AI field will gradually enter this state. ACT, Action Chunking, and Action Tokenizer include several original methods proposed specifically for robots. However, once a technology enters deployment, it ultimately becomes an engineering problem. To solve issues, you must continuously try new tools, regardless of whether they are sufficiently "fancy."

42HOW Robotics: So you didn’t propose Physical RSI first and then force your product into that concept?

Zhong Yuan Xin: No. We weren’t aiming for Physical RSI from the start.

We started with Reward, found it valuable, and decided to extend it. After deploying with Yuantu, we realized the upper-layer VLM should be handled by a foundation model, while we provide the Skill and Harness. Reward can evaluate; after Policy deployment generates new corner cases, those data points optimize both evaluation and strategy. As these pieces gradually came together, we realized they formed what is essentially RSI.

It was through doing the work that we discovered this path ahead.

42HOW Robotics: Many people are talking about RSI now, but their definitions differ. How do you understand it?

Zhong Yuan Xin: RSI has many levels. Letting AI conduct research and write papers can be called RSI; letting AI write operators and optimize Harness can also be called RSI. The distant endgame is AI designing its own architecture, data sources, and training methods.

In my view, the more realistic path today is integrating models with Harness. On one hand, the model improves continuously through data and results; on the other, the Harness evolves alongside model iterations. The same applies to embodied AI. How to evaluate, and how to make evaluation better, is the starting point of the closed loop. More accurate evaluation leads to better Policy. Once deployed, Policy generates more advanced corner cases, which in turn optimize evaluation.

Physical RSI enters the workstation before humans leave

42HOW Robotics: I just watched your demo on the Yuantu production line. The same station swapped different control boards, but there were still staff members assisting on-site. It hasn’t achieved full autonomy yet, correct?

Zhong Yuan Xin: Not fully autonomous yet, so we are still some distance from the ultimate goal of RSI.

If the Reward Model is good enough, the first step can at least replace human judgment in determining whether an action succeeded or failed. Achieving this enables a low-efficiency form of unattended operation: detect failure, terminate the task, and retry.

Once another model takes over after a failure to complete the motion, the entire closed loop truly begins to operate. Only then can we train and iterate more rapidly.

42HOW Robotics: On-site data collection and training took only about ten minutes. Have you stopped treating zero-shot success rate as the most important metric?

Zhong Yuan Xin: We don’t say we don’t pursue it. Initial success rate remains important because it determines subsequent training efficiency.

Failures come in many forms. Missing by one meter is a failure; missing by just one millimeter at the end is also a failure. When the model is poor, the resulting failures may involve foolish motions, and such data doesn’t need to be abundant. What’s truly valuable are failures that occur when the success rate is already high—missing by just one millimeter or one centimeter. These failures are rare but closer to the capability boundary.

You shouldn't deliberately make robots fail just to collect failure data. You should work like a smart person: even smart people make mistakes, but those mistakes are more valuable. This is the same logic as corner cases in autonomous driving: what was a corner case five years ago may no longer be one today. The most valuable data always moves along the capability frontier.

42HOW Robotics: But autonomous driving accumulates corner cases by relying on large-scale fleets. Robots are deployed one workstation at a time, so failure data remains very sparse.

Zhong Yuan Xin: It certainly won't happen as fast as in autonomous driving, but each piece of data may be more valuable. There are no shortcuts in industrial scenarios; you must deploy on-site to acquire it. Companies shouldn't expect robots to accumulate cases as quickly as autonomous vehicles do, but every company will eventually have to take this step.

42HOW Robotics: In the coming months, how will you determine whether Physical RSI is a genuine method or just a concept that holds up temporarily?

Zhong Yuan Xin: First, we need to verify whether it enables the system to 'learn better' or 'learn faster.' We have conducted preliminary validation, but it is not yet sufficient to establish it as a solid strategy.

The first step is to let the Critic automatically judge whether human intervention is needed; the next step is to have the model generate the intervention actions; further on, when facing a new skill, you can place the robot directly into the environment to explore, allowing it to learn the task on its own. That is the ideal state.

42HOW Robotics: During this process, which capabilities should come from the base model, and which must be learned on-site?

Zhong Yuan Xin: There is no need to learn all the underlying physical laws on-site, such as friction, gravity direction, object weight, and the differences between metals, rubber, and wood.

However, when it comes to a specific PCB board, determining the optimal center of gravity for grasping and avoiding damage to components above are knowledge points that base models rarely provide directly. These require robots to adjust within specific scenarios. Underlying physical laws come from pre-training, while capabilities for specific materials and tooling come from deployment sites.

42HOW Robotics: What standard must be met before you consider that the closed loop of Physical RSI has begun to function?

Zhong Yuan Xin: If we assume ten tasks have been trained, and by the eleventh task, the post-training cycle is significantly shorter than it is now, that would indicate the closed loop is starting to work.

I believe this milestone should appear before deployment reaches 100 units.

42HOW Robotics: Why decouple this from the 100-unit deployment? Doesn't more deployment lead to easier data accumulation?

Zhong Yuan Xin: A hundred units could simply be the same workstation with a few types of actions replicated across 100 devices; in that case, RSI is not necessarily required. RSI becomes more critical only when each of the 100 units corresponds to many different workstations.

According to our current plan, we will not cover a very wide variety of workstations next year. Instead, we will deploy a certain quantity per workstation type, reaching a total of 100 units. These two metrics cannot be conflated.

42HOW Robotics: Why is Yuantu willing to collaborate with you on this? Was their primary pain point the deployment time, or a shortage of labor?

Zhong Yuan Xin: Both. Take the example of swapping circuit boards for testing: human workers require almost no training and can adapt quickly. In contrast, debugging traditional robotic arms or industrial automation solutions can take half a day, tying up production line time.

They also face significant difficulty in recruiting staff. Beyond the large-board testing we saw today, there are other processes involving different boards and server testing, some of which take place in poor working environments. As capacity expands continuously, it becomes hard to hire and retain workers, creating strong motivation to replace manual labor.

42HOW Robotics: Many robotics companies avoid industrial scenarios because customer requirements are stringent and tolerance for error is low. As a later entrant, why did you still choose industry?

Zhong Yuan Xin: I do not believe the industry has reached a consensus that "industrial is the hardest, while other scenarios are easier." Each company simply has different partners.

We started from algorithmic needs. Relying solely on imitation learning with massive amounts of ego-motion data or teleoperation data to achieve long-term high success rates in household tasks like washing dishes or wiping plates—I have not seen complete validation of this, nor do I have sufficient confidence. However, by introducing reinforcement learning and post-training, certain tasks can be deployed first. We have already observed at Yuantu that post-training can further push success rates upward.

Post-training requires an environment with clear evaluation metrics and well-defined boundaries between success and failure. Following this criterion, industrial scenarios remain the most suitable.

42HOW Robotics: Another approach is to enter more inclusive scenarios first, where high success rates are not strictly required. For example, commercial services, human-robot collaboration, or specific household tasks.

Zhong Yuan Xin: I agree with the value of inclusive scenarios. Logistics sorting possesses both clear evaluation criteria and a degree of inclusivity, making it indeed a quite suitable scenario.

42HOW Robotics: So why don't you do sorting?

Zhong Yuan Xin: The sorting sector is too competitive. (Laughs)

For latecomer companies, there is no need to repeat problems that many teams are already working on and rapidly delivering. We focus more on flexible production with small batches and multiple runs, as well as tasks that traditional automation finds difficult to solve, such as handling non-rigid objects like wire harnesses.

If a production line requires no flexibility, traditional robotic arms can handle it, making embodied AI unnecessary.

42HOW Robotics: Isn't the tolerance for failure higher in home scenarios?

Zhong Yuan Xin: The home environment is not necessarily simple. When targeting consumers, you must also consider returns, after-sales service, and user satisfaction. Issues with large robots often require local engineers to handle them. Both industrial and domestic applications have their own difficulties; I do not believe that industrial applications are impossible, nor do I think the home sector is necessarily easier.

Our endgame remains robots applicable to more environments, and we will also consider the home. But at this stage, we start with industry because it is better suited for validating Physical RSI. I cannot yet determine whether it will become the optimal paradigm for home scenarios in the future.

42HOW Robotics: You have publicly reported a POC success rate exceeding 99%. What is still missing to move from a single POC to production line integration?

Zhong Yuan Xin: A POC is essentially completing an MVP. In the scaling phase, beyond success rate, stability is critical. Before entering the production line, extensive stress testing is required to ensure both hardware and models are stable. Additionally, the system must be integrated with factory workflows to establish safety fallbacks and enable parallel operation.

A target of one hundred units next year signals the start of scaling efforts, but this number alone does not prove cross-station generalization, nor does it confirm that the RSI closed-loop is established.

"Algorithms Are Not a Moat"

42HOW Robotics: Physical RSI imposes high engineering demands. Your team cannot consist solely of algorithm engineers, right?

Zhong Yuan Xin: Correct. Currently, more than half of our staff are engaged in engineering work.

We are not a pure algorithm company; we focus more on overall iteration efficiency. If the robot body is unstable or has minor issues, it severely impacts training and deployment.

42HOW Robotics: If the core lies in the engineering closed-loop rather than an unreproducible algorithm, what is RiMian's moat? Could other companies with similar resources achieve the same result?

Zhong Yuan Xin: That’s essentially it.

We made a clear decision from the start: our moat consists of only two elements—data and infrastructure. Algorithms are not a moat. If an engineer who masters algorithms moves to another company, their algorithmic capabilities go with them; once a paper-based method is published, academia and peers can replicate it quickly.

What cannot be taken away is data and integrated infrastructure. Infrastructure includes cloud training systems, data pipelines, and also the robot hardware itself. The hardware body is part of that infrastructure. Companies with stable hardware bodies can iterate models and adapt to new markets faster.

42HOW Robotics: You mentioned that Physical RSI does more than improve algorithms; in the future, it might also improve end-effectors. How would that approach work?

Zhong Yuan Xin: This remains an early-stage concept that has not yet been validated on a large scale.

After running experiments in many industrial scenarios, we found that there is no universal end-effector capable of covering all tasks or performing efficiently across every task. End-effectors will ultimately require some degree of customization. The question is whether foundation models can adapt after customization and whether delivery costs remain manageable.

Take cable plugging as an example. In demos, you might plug in just one cable; but on a customer’s device, ten cables may sit side by side. The end-effector must first be able to reach inside. We cannot have engineers design solutions from scratch for every scenario. Ideally, the system should first check if a ready-made end-effector exists; if not, it designs or selects a new one. The algorithm attempts it, switches if performance is poor, continues training, and repeats until performance is sufficient.

Only by automating this loop can customized end-effectors achieve scale. But we are still far from achieving that.

42HOW Robotics: In fine-manipulation scenarios like those at Yuantu, do you opt for dexterous hands or grippers?

Zhong Yuan Xin: We have tested many dexterous hands, but very few products are truly capable of serious work. Low-DOF dexterous hands are not fundamentally different from a large gripper, while high-DOF hands struggle to operate stably over the long term. Therefore, in industrial scenarios, our current priority is to perfect the gripper.

However, tactile and force feedback remain essential; even grippers can incorporate sensors. Fine-manipulation scenarios involving servers demand high levels of force feedback, which will influence our future data collection methods.

42HOW Robotics: Tactile sensing has become a new direction for many teams. What is the biggest bottleneck right now: hardware or models?

Zhong Yuan Xin: It is both, and the two may be coupled.

How tactile and force capabilities generalize, and how models learn across different sensors, remains largely unexplored. A human seeing pressure values can roughly infer the state of an object; today’s models still rely primarily on vision, such as checking whether an object has dropped. But knowing an object hasn’t dropped does not mean the model understands whether the applied force is large or small.

It could be that tactile hardware lacks standardization, preventing software from scaling; or it could be that immature model architectures make teams hesitant to deploy tactile hardware at scale. We have attempted to integrate vision, tactile, and proprioceptive states into Ego representations, but the results of fused training currently underperform compared to keeping separate branches.

42HOW Robotics: Reward Models can mitigate some proprioceptive differences, but within your current framework, how difficult is cross-embodiment adaptation really?

Zhong Yuan Xin: It remains difficult. The Reward Model only reduces part of the difficulty. To achieve cross-embodiment, you first need cross-embodiment data, which is inherently hard to acquire.

42HOW Robotics: Is cross-embodiment a long-term goal for you? The demo today featured dual robotic arms and did not showcase mobility. Do workstations in Far Sight’s deployments require movement?

Zhong Yuan Xin: Mobility is still required. Mobile chassis will be added soon.

42HOW Robotics: So it was simply not demonstrated this time?

Zhong Yuan Xin: Correct. We did not demonstrate mobility this time. However, the robots deployed in next year’s batch of 100 units will definitely include chassis.

42HOW Robotics: At what stage do you plan to validate cross-embodiment?

Zhong Yuan Xin: We must start from customer needs. When existing embodiments cannot fulfill tasks and customers genuinely require a different embodiment, we must address it. Under the current paradigm, while cross-embodibility is methodologically feasible, it does not mean that data and delivery are ready.

42HOW Robotics: Do you think deployment teams will eventually reach a consensus to adopt Physical RSI?

Zhong Yuan Xin: Physical RSI is a tool provided to everyone.

If this tool works well enough, deployment engineers in the future will only need to monitor whether the final result is correct, using rewards for training and adjustment. Without Physical RSI, they would be like the demonstrators on-site today, constantly performing manual teaching. Essentially, this comes down to how much effort and time a deployment engineer must invest for a single delivery.

42HOW Robotics: As fewer people are needed for on-site deployment, what will the role of engineers become?

Zhong Yuan Xin: Humans will still be needed initially, but the number of personnel required for deployment will gradually decrease. This precisely means that more scenarios can be deployed: previously, one person might have spent a month deploying a single production line; in the future, one person could deploy ten or even a hundred lines in a month.

This does not mean human demand will disappear. If we truly achieve unmanned deployment or unmanned training of new skills and delivery further down the line, engineers can shift their focus to infrastructure itself or other higher-value problems.

42HOW Robotics: What do you think is the most critical problem for the industry to solve at this stage?

Zhong Yuan Xin: The most macro-level short-term issue is how to integrate physical-native modalities into models. There is still a lack of mature solutions for truly fusing vision, force, and touch and scaling them up.

The long-term issue is how to balance scaling with customization. If systems span different robot bodies, more data must be trained, making scaling difficult; if they do not span different bodies, many customer tasks cannot be completed. Whether automation design can reduce modifications to the robot body while retaining sufficient scenario adaptability is a significant challenge.

42HOW Robotics: You haven’t purchased data at the same scale as some other companies, have you?

Zhong Yuan Xin: We’ve hit a few snags. We bought some data but quickly found that the data paradigm and delivery methods didn’t align.

If the delivery relies heavily on ego data, then you should purchase ego data; if it depends on UMI plus a small amount of real-robot data, then you should purchase UMI. If the company hasn’t clarified its delivery paradigm, the data purchased may be useless. That’s why we haven’t committed to large-scale data procurement.

Now, we want to explore non-invasive, ambient collection: gathering data naturally generated during operators’ work without disrupting factory production. Operators could wear head-mounted devices, but we aim to avoid adding anything to their hands. The core goal isn’t just eliminating occlusion, but ensuring no interference with operations. This requires co-designing hardware and algorithms, and it is still in the research phase.

42HOW Robotics: Many robotics companies now hire data leads from autonomous driving. Between autonomous driving teams and large language model (LLM) teams, who is better equipped to build an embodied AI data closed loop?

Zhong Yuan Xin: Each has its strengths.

Autonomous driving data pipelines have handled massive volumes of visual and temporal data from the start, making them more adept at multimodal data storage, format conversion, frame alignment, and time synchronization. LLM data formats are relatively simpler—historically focused on image-text pairs—but these teams typically excel at deeper data cleaning, deduplication, and quality control.

Embodied AI data requires both capabilities. It is multimodal and temporal, yet also faces challenges regarding data quality, redundancy, and ratio balancing. You can’t simply transplant an autonomous driving data pipeline, nor can you rely solely on the cleaning methods used for language models.

42HOW Robotics: How much data was used in total at the event today?

Zhong Yuan Xin: This wasn’t a dedicated data collection effort. More accurately, it was an end-to-end demonstration of data acquisition, training, and inference on-site. The data volume was small—only about ten minutes’ worth.

If deploying to a new production line where tasks differ significantly from the model’s current capabilities, additional data collection might be needed for several hours. That’s roughly the scale we’re talking about.

42HOW Robotics: What drove your transition from autonomous driving to robotics? Did you feel the challenges in autonomous driving were largely solved?

Zhong Yuan Xin: No. I believe autonomous driving is still at Level 3, with a considerable distance remaining before true Level 4 autonomy is achieved.

My primary reason for leaving is that few are exploring new paradigms in this field anymore. Conducting experiments in autonomous driving now requires massive amounts of data and compute power, making it increasingly difficult for university labs to participate. As universities drop out, industry discussions around model architecture have diminished, with most efforts focused on solving engineering problems. I wanted to continue exploring new directions and questions, so I turned to robotics.

Robotics still allows individual universities or PhD students to build valuable models with relatively limited resources. Data, representations, and delivery paradigms have not yet converged. This presents difficulties for the industry but also opportunities for researchers.

42HOW Robotics: As a latecomer company, are you optimistic or pessimistic about the industry?

Zhong Yuan Xin: I am optimistic about delivery. Robots can create tangible value and enter production lines.

However, I am relatively pessimistic about how far robots are from AGI, or from becoming truly general-purpose machines that can do anything. The release of new technologies does not make me more pessimistic; it simply provides us with more tools to use.

What we can do is avoid maintaining systems with heavy reliance on rules from the start, and refrain from blindly purchasing data before the data paradigm has converged. First, close the loop among models, data, hardware, and deployment in real-world scenarios, and observe which parts genuinely accumulate capabilities through continuous use.

42HOW Robotics: You have previously emphasized WPM and WRM. What specific role does Real-world RL play within the Physical RSI closed loop?

Zhong Yuan Xin: Using reinforcement learning for post-training involves a constant trade-off between real-robot and simulation environments. Real robots can handle fine-grained operations because the physical dimensions of the real environment are accurate; however, training on real robots is inefficient, requiring at least one robot and a dedicated real-world setup. To double efficiency, you must add another robot and workstation, which significantly increases costs.

The challenge with simulation lies in replicating the site with sufficient accuracy. This dynamic changed after the release of GPT-6—though this applies to other models as well, with attempts using code for simulation reconstruction already emerging in July. By providing a photo of the site along with minimal additional information, the model can reconstruct a simulation environment with reasonable accuracy. If we calibrate key dimensions more precisely, we can shift part of the learning process into simulation.

Previously, relying entirely on real-robot post-training meant that each delivery might require occupying multiple complete units. Our current vision is: the real site provides tasks, errors, and feedback; simulation amplifies training efficiency; and results return to the real robot for validation. Only then is Real-world RL likely to become a scalable delivery process.

Whether Physical RSI holds true should ultimately not be judged by the popularity of its name. A more direct test is whether engineers who currently stand by workstations can leave earlier next time; and whether, after a system completes ten tasks, the eleventh task can be deployed with less data and less time, without compromising safety and stability.

If these metrics do not show sustained improvement, so-called recursive evolution remains merely more complex on-site debugging. Only if they begin to stabilize and decline with the number of deployments will the workstation at Ri Mian Kai Wu and Yuan Tu Wei Lai, which currently still requires manual assistance, become the place where the true physical intelligence loop begins to turn.