Exclusive Interview with Zhou Shunbo: We Can Build a Robot People Love, Even Without AGI

In founding Oula Wanxiang, Zhou Shunbo had already been working on robots for nearly 12 years.

In 2015, he joined the laboratory of Academician Yangsheng Xu at The Chinese University of Hong Kong, which marked the true beginning of his work in robotics. He participated in the development of a panda-shaped therapeutic robot, which was later deployed at Sha Tin Hospital in Hong Kong to accompany children with autism and elderly patients with Alzheimer's disease. At that time, such robots were still primitive; their tactile sensors had to be manually attached one by one, far from being mature or reproducible products. Yet, Zhou Shunbo had already formed a clear aspiration: to move robots beyond the laboratory and turn his favorite technologies into products that ordinary people could genuinely use.

In 2016, he pursued his doctoral degree under Professor Liu Yunhui. During the interview, when asked why he wanted to pursue a PhD, he replied, 'To better prepare for entrepreneurship.'

From visual positioning and state estimation to the dynamics and control of industrial vehicles, Zhou Shunbo researched how to enable real-world machines to operate stably in environments lacking GPS and experiencing constantly changing loads. His first-author paper on navigation and control for industrial tractors received an honorable mention for the IEEE RA-L Best Paper Award.

But as his research deepened, his enthusiasm gradually cooled. Learning physics, mathematics, control theory, programming, and hardware meant mastering so much, yet it only allowed robots to barely function in limited scenarios. The vision of a robot product that could be owned, used, and even personally created by ordinary people seemed to grow increasingly distant.

After completing his PhD in 2020, Zhou Shunbo spent nearly a year with his mentor preparing for a startup focused on underground mining robots, aiming to use mobile robots for material handling tasks in mines. The project eventually halted due to resource constraints and corporate governance issues. This was his first realization that knowing how to build a robot does not equate to understanding product management, delivery, maintenance, and commercialization.

In 2021, Zhou Shunbo joined Huawei through its 'Genius Youth' program, becoming the only candidate selected thus far for a topic in intelligent robotics. He built Huawei’s Embodied AI team from scratch, later serving as the head of Huawei Cloud’s Physical Intelligence Innovation Lab and as a company-level chief expert. The team grew to over a hundred members, bringing robotic technology out of academic papers and proof-of-concept (POC) stages and into real-world industrial projects.

There, he witnessed another layer of difficulty faced by robots once they leave the laboratory. During POC phases, factories welcome new technologies enthusiastically; however, once truly integrated into production environments, they immediately become demanding, questioning reliability metrics like five nines uptime, production cycle times, single-task success rates, and ROI. Academic research revealed how difficult it is to build robots, while industrial deployment showed him just how far there is between building something and having it genuinely adopted.

The thing that truly changed his judgment was an unassuming case within Huawei.

A product manager who rarely wrote code and had no robotics experience independently developed a flexible spool picking application using the team's robot development platform. By previous standards, this would have taken at least one professional developer a month.

For the first time, Zhou Shunbo realized that while general-purpose robots might still be distant for ordinary users, the barrier to developing robot applications was beginning to lower. With new models and tools, even those without a professional background in robotics could develop new tasks for robots through demonstration.

'At that moment, many things clicked into place.'

The product he had thought about for years, and once believed was impossible to achieve, suddenly presented a viable path to implementation.

Zhou Shunbo decided to start his own business. This year, the robotics industry also began to reward another group of people who had persisted for over a decade. Wang Xingxing, in his 17th year of working on robots, brought Unitree Robotics to the eve of its listing on the STAR Market; UBTECH, Dobot, Yunji, and other companies had already entered capital markets, with more robotics companies queuing to submit their registration statements. At a time when pioneers were starting to reap rewards and newcomers were flooding in, he felt he hadn't arrived too late—the timing was actually just right.

In March 2026, Zhou Shunbo left Huawei to found Oulaxiangwan. This time, he did not return to his most familiar industrial scenarios but chose to enter the home sector.

Not because homes were easier, but because Zhou Shunbo believes data has primacy: if the endpoint of robots is real life, the data used to train them should come as much as possible from places where people truly live. Factories, warehouses, and hotels can generate large amounts of data, but they are prone to converging into customized solutions under the pressure of delivery and ROI; only home and home-like scenarios can continuously generate sufficiently diverse and long-sequence human-robot interaction and physical feedback.

Euler Universe aims not to build the currently hyped bipedal humanoid, but a wheeled home robot. It won't be ready-to-use out of the box; instead, it can be taught, raised, and have its skills expanded by users. Starting with organizing floors, sofas, coffee tables, and desktops, it will gradually evolve based on data from real households. Its earliest target audience isn't ordinary consumers lacking patience for robots, but geeks similar to early Bambu Lab users—people willing to tinker, allow product imperfections, and participate in creation.

Liu Qin, founding partner of Sequoia Capital China, describes Euler Universe as "Physical OpenClaw": rather than striving alone for an embodied AI "DeepSeek moment," the company transforms continuously improving model capabilities into robot skills that users can utilize and develop.

Although the product has not officially launched, this approach has already attracted capital market attention. Within three months of its establishment, Euler Universe completed three consecutive funding rounds, with one round in June reaching hundreds of millions of yuan. Investors include Hillhouse Venture Capital, Sequoia Capital China, China Merchants Group Venture Capital, Baidu Venture Capital, among others. The first prototype units are expected to reach early users' homes by year-end.

Zhou Shunbo believes in the future of embodied AI but refuses to leave the viability of a product entirely dependent on a technological endpoint that hasn't yet arrived.

"Even without AGI," he said, "we can still make a robot that people like."

The following is an edited conversation between 42 Wave and Zhou Shunbo:

I. From Platform to Product: Why Home?

42 Wave: Briefly introduce Euler Universe. What exactly are you trying to achieve?

Zhou Shunbo: Oula Wanxiang is a home-based Embodied AI company. We focus on two main objectives: first, democratizing robot technology to empower every enthusiast, enabling more people to develop and nurture their own robots; second, extending access to certain expensive and scarce professional services currently available in human society to every household.

The former dictates that we will start with the Maker and developer communities, while the latter determines that our ultimate goal lies within the home environment. We aim to lower two types of barriers simultaneously: the threshold for creating robotic capabilities, and the threshold for ordinary households to obtain real-world services.

42 Hertz: Why did you leave Huawei to start your own business?

Zhou Shunbo: I have always been passionate about robotics and hoped to turn this passion into products accessible to everyone around me. My goal has remained unchanged throughout these years. In fact, choosing to pursue a doctoral degree was also driven by this ambition.

However, during my PhD studies, my enthusiasm gradually waned. I pursued my doctorate from 2016 to 2020. At that time, building a robot required expertise in physics, mathematics, control theory, programming, and hardware engineering. After mastering all these disciplines, one could barely develop a robot capable of moving only in predefined scenarios. The barrier to entry was simply too high, making it impossible to bring robots to everyone, which led to increasing frustration.

Over the past few years, however, I witnessed a paradigm shift in industry technology while at Huawei. Today, developing robots no longer demands highly complex modeling; instead, the field has gradually shifted toward data-driven approaches. With sufficient data, computing power, and an effective pipeline, robots can now be trained. This revealed the possibility of bringing robots to the mass market.

Another thing that deeply resonated with me was during my time at Huawei, where we worked on a To-B platform aimed at lowering the barrier for developing industrial robot solutions. One day, one of our product managers—who rarely wrote code and had no prior experience with robots—used this platform to independently develop a flexible coil picking application. Under previous methods, a professional developer would have needed at least a month.

In that moment, I realized: even if I couldn't directly sell a general-purpose robot or model to users at that time, wouldn't everyone be able to play with and utilize robots through such tools? That instant connected many dots.

So when you've been thinking about a goal for years, wanted to pursue it in the past but couldn't, and now the possibility of achieving it has finally emerged, I feel I would definitely choose to jump in and try.

42 Hertz: But you were already leading a team of around 100 people at Huawei, and had built the platform from 0 to 1. Why didn't you stay at Huawei to continue working on robots?

Zhou Shunbo: We proved that this toolchain can indeed significantly lower the barrier to robot development. In the next stage, I hope to find a high-value scenario to dive deep into, and truly run through the data closed loop, product closed loop, and commercial closed loop simultaneously. Only by getting the flywheel spinning can I identify where the bottlenecks in the platform are and determine what needs to be addressed next.

From Huawei's perspective, a more logical choice would be to go from 1 to N, continuing to replicate the platform and keep "selling shovels," rather than diving deeply into a single vertical scenario themselves.

42 Hertz: You don't want to continue selling shovels?

Zhou Shunbo: Correct. My own judgment is that Embodied AI's so-called "general" capabilities aren't actually general yet. Especially with platforms, you must spin up the data flywheel within a real-world scenario, make the numbers work, and then gradually expand outward based on that scenario. Otherwise, if you keep selling shovels but no one is truly using them to dig anything up, the platform won't know how to evolve.

If I stayed at Huawei, I wouldn't be able to achieve this kind of vertical integration. Entrepreneurship gives me decision-making power at least. Current investors are investing because of what we aim to do; our goals are relatively aligned, and no one is demanding that I pivot to another type of company. Therefore, I can proceed according to my own judgment.

42 Hertz: When you were at Huawei, did you experience FOMO seeing so many startup projects outside?

Zhou Shunbo: Mostly no. Because I’ve been in the industry for a long time, so I roughly know every development and how many things are actually built.

When does FOMO arise? It happens when someone suddenly achieves something you don’t understand how they did it. That’s when trouble starts.

I remember 1X once caused anxiety among our team members. They announced plans to put robots into users’ homes at a price of $20,000. Our co-founder Zhang Jing was particularly anxious that day: almost the entire supply chain for this robot is based in the US. If you break down the BOM (Bill of Materials), it likely costs more than $20,000. So why would he be willing to sell at a loss? Did the model already achieve some massive breakthrough, only needing to push the robot into households to collect data and conduct evaluations? Was the robot already at the stage of the Internet’s "hundred-regiment battle"?

**42 Hertz:**So it wasn’t that its performance was exceptionally strong, but rather that it was sold too cheaply, which triggered your FOMO.

Zhou Shunbo: Exactly. Sold cheaply, and certainly at a loss. You’d think, why would he be willing to lose money on this? Is it just one step away from success? That’s what I mean by FOMO: an event appears that is difficult to deduce using conventional technological logic.

**42 Hertz:**How did you eventually rationalize it?

Zhou Shunbo: Later, from a technical perspective, we found that everyone was still far behind.

Moreover, the capital market is often honest. If 1X had truly demonstrated generational-leading capabilities, we might not have known immediately, but investors certainly would; if investors saw it, the company’s financing and valuation would reflect that. This was one of the side references we used to calibrate our judgment at the time. Of course, valuation does not equal technology; it can only serve as corroborating evidence.

"I didn't wait for the track to become hot enough before entering, nor do I get anxious just because others launched first or raised funds first. What truly changes my judgment is when I see a technical fact that was previously unexplained and later proven valid."

42 Hertz: You've worked in industrial scenarios at Huawei for so long. Why are you now making home robots? Aren't homes more difficult?

Zhou Shunbo: I hold some unconventional views. On one hand, these views come from our technological investments over the past few years; on the other hand, if we look toward the endgame, I believe data is primary.

Of course, if you want to train an embodied model in its ultimate form, what kind of data is needed? This is hard to define, and no one can clearly define it right now. But some broad directions are gradually converging. For example, data needs sufficient diversity, must include many long-sequence tasks, and also need to encompass real human-robot interaction and physical world feedback.

If you believe in the primacy of data, you'll find that home and home-like scenarios are the best places to accumulate this data, without exception. The next question then becomes: Can you find Product-Market Fit (PMF) in home scenarios? But that's another issue.

Of course, there's another reason I chose the home sector. When I was at Huawei, I was essentially given a 'semi-open prompt' to work on industrial-leaning scenarios. This was determined by the nature of Huawei as a company. However, this experience actually gave me some negative incentives. To me, industrial and commercial scenarios are far from as simple as people imagine.

At that time, we served both Huawei's own factories and manufacturing departments, handling packaging, material transport, and picking for 3C products like mobile phones, as well as serving various enterprise clients.

You'll find that during POCs, everyone is happy: it's a new technology, and I can use this new technology to run through the tasks. But even within Huawei's own factories, once you say the technology has been validated and hope to enter the formal production environment, the factory immediately gets serious and tells you: I need five-nines reliability; what is my production cycle time; what is the success rate per task; how should ROI be calculated?

"If you truly want to solve a high-value problem in a factory, I actually feel that it's not necessary to focus on Embodied AI. You can use mature automation technology combined with some AI to create an AI Plus solution. The ROI is likely calculable, but it wouldn't be as 'sexy,' nor would it necessarily be an AI-native product."

"I already knew how industrial solutions should be done during my time at Huawei. Oula Wanxiang didn't choose the home sector as a second-best option because industrial applications were unworkable. Quite the opposite: having seen that path, I became even more certain that I wanted to build a truly data-driven product capable of continuous evolution."

42 Hertz Radio: There are so many household tasks and non-standard environments. Why do you think it's not as difficult as people imagine?

Zhou Shunbo: If you expect the robot to do everything, then yes, households are very difficult. But we will find the intersection across three dimensions: First, is this a high-frequency need? Second, are the loss and safety costs of failure sufficiently low? Third, is the current technology achievable?

"For example, picking up toys from the floor. Even if the policy has a single-attempt success rate of only 70%, if the first attempt fails, I can try again. Once the child stops playing, there is no strict rhythm for the tidying task, nor is there a requirement for one-time success. The probability of failing three times in a row is already quite low and does not necessarily affect user experience."

"Take cleaning pet hair off a sofa as another example. Even when humans do it, they brush from start to finish with a brush; they don't use a spotlight afterward to check for remaining hairs. There is no absolute 'done' or 'undone'; as long as the action is taken, there is an effect."

"There are actually many such tasks at home. Home network conditions are usually good, allowing for edge-cloud collaboration on capabilities that don't require high real-time performance, further reducing the computing power costs of the main body. Most importantly, finding an elegant product definition: it must be valuable to users while being achievable with today's technology."

II. Defining Robots in Reverse Based on Human Data

42 Hertz:** Robot vacuum companies are already trying to give their products a hand, starting by picking up items on the floor. Although the first generation’s performance was mediocre, it had existing users and supply chains. Why can’t home robots emerge incrementally along the lines of robot vacuums?**

Zhou Shunbo: I don’t believe this kind of evolution can naturally grow out from the extension line of any existing product category.

If a robot vacuum grows a hand and uses traditional rule-based hard coding, it remains a previous-generation robot; if it uses data-driven approaches, you need to remote-control that robot vacuum to collect "picking up" data across various households. Assuming it eventually picks up items on the floor well enough, in the next generation, I’d want to raise its waist height so it can reach sofas and tables—it has essentially become another type of robot altogether. You’d then have to re-converge the hardware platform, re-collect data, and retrain models.

The biggest problem isn’t that the first generation’s capabilities were weak, but rather that what was learned previously is difficult to continuously transfer to next-generation products. Therefore, our choice is to first converge data collection onto humans, rather than onto a specific generation of robots. Humans pick up toys on the floor, fold clothes on sofas, and organize items on desks—all actions performed by human hands. By accumulating these human-centric data first, future adjustments to the hardware platform will still allow the data to be migrated effectively.

Of course, humans aren’t robots, which creates a gap. How to perform retargeting and accurately estimate spatial states are issues we’ve been working on resolving over many years.

42 Hertz:** You repeatedly emphasize that data comes first. But if data is most important, intuitively, you should start a company focused solely on the "brain" for robots. Why did you ultimately choose to develop an integrated software-and-hardware product instead?**

Zhou Shunbo: Prioritizing data doesn’t mean we won’t build models. On the contrary, models are a crucial component within the entire data pipeline.

Oula Wanxiang currently invests in three core technologies: first, the toolchain for closed-loop iteration from data to models; second, foundational models; third, Agent OS. These three elements work synergistically and none can be omitted. Models are important, but they do not constitute the entirety of our company.

The 'data first principle' I understand primarily reflects the importance we place on data and how it guides our internal decision-making. The robot we ultimately aim to build must be perfectly matched with our entire data pipeline and data collection methods, while providing users with the smoothest possible experience. In a sense, it is a robot defined by this data-first principle.

42 Wave: How specifically does the 'data first principle' determine the form of the robot?

Zhou Shunbo: Let me give a very practical example. Once a robot is sold into a user's home, there will inevitably be tasks it cannot perform well. A direct interaction method is for the user to guide the robot's hand through the motion, much like an adult teaching a child.

Therefore, when designing the hardware body, you must consider factors such as chassis size, robot height, and workspace dimensions to ensure that guiding it feels comfortable for the user. Only if the interaction is natural enough can users quickly generate rollout data, allowing that data to feed back into the system.

We favor dual arms because humans naturally operate with two arms. Data collected from human hands can be transferred to single-arm or dual-arm robots. Conversely, whether the base is wheeled or bipedal may not be as critical for the upper-level model. What truly needs to be captured in human data is how the left hand moves, how the right hand moves, and how the human body as a whole moves through space. As long as the low-level motion controller is sufficiently robust, the upper-level model can provide spatial targets for the next step, which both wheeled and bipedal bases can execute.

42 Wave: So do you believe a single arm is insufficient? Many storage tasks can be completed with one hand.

Zhou Shunbo: A single arm can certainly handle many things, and we may release a single-arm version in the future. However, if we want the robot to continuously evolve alongside human data, dual arms are more natural because a large portion of human operations are performed with both hands.

We will produce as few robot categories as possible, unlike the automotive industry which offers many models. Our core product will undergo continuous iteration, gradually increasing load capacity, reliability, and capabilities. If a single-arm version could significantly lower the threshold for users to experience the robot and for developers to master skills, we would consider it. However, regarding whether to include autonomous mobile chassis, there is currently internal disagreement, and further validation is needed.

42 Radio: Why did Euler万象 choose a mobile chassis with dual arms? Has this form factor converged?

Zhou Shunbo: The term 'converged' is too broad; we need to be more rigorous technically. I can only say that it covers the main markets and tasks we have currently defined.

For the next three to five years, I will only consider wheeled robots. As someone from the robot hardware background, I remain deeply respectful of hardware limitations. Bipedal humanoid robots currently serve no purpose in homes other than dancing. They face significant issues regarding safety, energy consumption, noise, and reliability—problems that may not necessarily be resolved within the next three to five years.

In our early stages, selling products was aimed at continuing to iterate on manipulation technologies. Since manipulation requires continuous improvement, the mobile platform must be as stable as possible. A wheeled chassis allows us to concentrate resources on arms, data, and tasks.

42 Radio: If home robots are too large, users might feel overwhelmed; if too small, their load capacity and functionality may be insufficient. How do you balance size, load capacity, and functionality? Can you reveal the approximate dimensions of your product?

Zhou Shunbo: We cannot disclose product details yet. However, this is an excellent question. When we talk about home robots, we are actually discussing not just technology, but a product. Therefore, I believe this endeavor requires a very clever, even genius, product definition to move forward. Every product has boundaries; you must use these product definitions and boundaries to manage user expectations—that is, robots are not omnipotent.

42 Radio: After users receive the robot, how exactly do they teach it? How many times must they teach it before it truly learns?

Zhou Shunbo: We will categorize tasks into two types.

The first category covers tasks that already fall within the capability list of foundation models. With current technological maturity, users typically need to collect between ten and dozens of data points, which corresponds to about an hour of work in the physical world. Through post-training, success rates can be pushed to around 90% or 95%.

For example, a robot might default to placing keys in a storage tray at the entrance, but you prefer hanging them on a pegboard. By wearing a first-person view device and using your bare hands to demonstrate moving the keys from the tray to the pegboard a few times, the robot can internalize this as your preference.

The second category involves tasks completely absent from foundation models. We interviewed a geek who wanted his robot to tap a wooden fish at home because he often works and holds meetings there, needing to keep calm. This is a需求 we cannot easily anticipate beforehand, so it likely does not exist in foundation models.

In such cases, we use our skill training toolchain to train a very small policy from scratch, which may also require only dozens of data points. While it lacks strong generalization capabilities, I believe this is unimportant because it is designed specifically to satisfy this user's personal hobby. The user invests some effort in exchange for significant emotional value, which is a worthwhile trade-off.

**42 Hertz: Can skills trained in one household generalize to others?

Zhou Shunbo: Some can, while others do not need to. It depends on task complexity and type.

For the vast majority of users, if they are not making money by developing skills nor deriving joy from sharing, many tasks do not need to generalize. Once a robot is sold into your home, it stays there for life. It is your family companion and does not need to know how things are done in other people's homes.

Generalization only becomes important when the community grows large enough that people develop robot skills akin to apps on the App Store. For instance, writing Spring Festival couplets or tapping a wooden fish could be downloaded by other users. However, 'where my keys should hang' is inherently tied to a specific household and only needs to be effective within that home.

42 Hertz: Who are your first batch of users? Do you require them to understand robotics?

Zhou Shunbo: We have currently accumulated about 2,000 potential users within domestic and international robotics communities and enthusiast groups. The experience officer recruitment drive received over 1,000 long-form questionnaires, from which we finally selected approximately 200 of the most relevant users. Our product manager team is conducting in-depth interviews with these 200 individuals one by one; it is a gradual screening process.

42 Hertz: How do you screen these 200 people?

Zhou Shunbo: The questionnaire is very long, and we are already grateful that anyone takes the time to complete it seriously. The questions include home square footage, whether it is a single-level apartment, if there are thresholds exceeding 3 centimeters, housework habits, what their most painful household chores are, and their acceptance level regarding privacy boundaries for robots.

We prioritize letting robots enter homes with higher composite scores. This is not primarily to judge whether the person is "good" or not, but because early-stage robot capabilities are relatively weak, they need to enter environments where they can operate more easily. For example, larger spaces, fewer clutter items, and no need to navigate between floors.

42 Hertz: Besides the hard conditions of the home environment, what level of robotics know-how do you require from these people? Are they ordinary consumer users or developer-type users?

Answer: Many of these people are front-line developers at major tech companies. They possess general development skills, like robots, but do not specialize in robotics. We call them "generalist developers." We actually received many more questionnaires; some respondents were wealthy individuals, and some were even our investors, among others. We will place these users slightly further back in the queue, as they are closer to ordinary enthusiasts.

Of course, some colleagues within the company also wish to have their family members participate, such as their parents. However, since they are relatively older, even products matured for ordinary enthusiasts might not yet meet their needs. Therefore, we can only temporarily set aside this type of user.

Our users are roughly divided into the following layers: The first layer consists of professional developers, for whom we complete the initial version ourselves. Then, we roll out the product to a broader group of general developers for testing.

42 Hertz:** Are your beta testers still open for registration?**

Zhou Shunbo: Yes, we are continuously collecting registrations. It essentially serves as a user pool; we never have too many users, only too few. The only difference lies in the order of shipment. We have an official WeChat account that currently hosts only one post—the call for experience officers—which allows direct registration via the account.

III. The Robot Brain Is Not Just a Battle of Model Architectures

42 Hertz:** Which do you believe is more important: algorithms or engineering capabilities?**

Zhou Shunbo: Engineering capabilities might be more important than algorithms. Startups are not merely about technical research, and I do not believe any single algorithm can form a long-term barrier because technology will eventually become democratized and fluid. Even models, which may seem elegant and sophisticated to outsiders, reveal themselves to be entirely problems of engineering and resources when you actually work on them. Often, the issue is simply a lack of GPUs.

It is currently difficult to prove that any specific architecture is inherently superior to others. We conducted relevant experiments before, benefiting from Huawei’s relatively abundant resources. For any model architecture, the only method to equally enhance all of them and lift all performance curves together is to improve data quality. The reverse is also true.

When scaling up significantly, what is needed is not an elegant architecture, but a simple, user-friendly, and scalable one. Everything else is engineering: how to scale, clean, and unify data formats; how to optimize infrastructure; and how to manage cross-GPU clusters. If training is interrupted midway, all the compute power invested beforehand is wasted. In that case, it would be better not to have those GPUs at all, right?

42 Hertz: Is it easier for a robot brain company to emerge from a robotics firm or from a large language model (LLM) company?

Zhou Shubo: I don't think LLM companies can give rise to such firms. Andy Zeng, who focuses on generalist approaches, has been working on robots and specifically on robot learning. Deepak Pathak from Skild AI and his team originally came from the CMU Robotics Institute. Sergey Levine from Physical Intelligence (PI) and others have always come from traditional robotics labs. These institutions are essentially the "Whampoa Military Academy" of this field. These individuals possess a strong sense of the robot hardware itself as well as a deep understanding of robot learning. Therefore, I believe this capability is difficult to cultivate within LLM-centric companies.

42 Hertz: An engineer told me that a year after the release of Physical Intelligence’s π0.5, no domestic model has truly surpassed it in comprehensive capabilities. Do you agree?

Zhou Shubo: I agree. Currently, many people are focused on ranking various benchmarks. Consequently, there is a common saying in the industry that π0.5 is the "King of Real-World Deployment": it may never win on benchmark leaderboards, but it has never lost when deployed on real hardware. This statement realistically reflects the situation. I think PI is truly an excellent company; in some respects, it helps bridge the gap between Chinese and American models. While China indeed holds advantages in supply chain, hardware platforms, and data, I believe we are still at least 6 to 9 months behind in terms of model architecture.

42 Hertz: Why is PI so strong?

Zhou Shubo: I believe this involves several dimensions.

First, they are indeed one of the most cutting-edge embodied AI model research institutions globally. The founders come from top-tier universities, and most talent in this global field is cultivated by these academic systems. Objectively speaking, they started earlier, conducted deeper research, and hold more advanced concepts.

Additionally, there is another crucial point: Although PI is a company, its expectations from the capital market differ significantly from those of domestic firms. It is not necessarily a "gap," but rather a difference in the expectations each side faces.

From an organizational perspective, PI feels more like a lab. The current environment at PI is extremely relaxed. Although it is a commercial company, there are actually no commercial KPIs. Such a lenient entrepreneurial environment is still hard to imagine in China today.

42 Hertz: So how do you view the debate between routes like VLA and WAM?

Zhou Shunbo: I think names aren't that important; what matters is the essence.

Embodied models must ultimately output actions. 'V' represents multimodal perception of the environment, 'L' represents human intent and preferences, and finally, 'A' is the output action. In this sense, VLA is already an endgame concept. WAM also needs to understand vision, intent, and output actions; essentially, these three things cannot be bypassed. How the black box in the middle is designed constitutes the model architecture. The different names people use reflect both technical differences and considerations for establishing labels and facilitating fundraising.

However, I remain cautious about WAM. It couples video generation with action, meaning the upper limit of action capabilities is constrained by video generation. But if a video model could perfectly understand physical laws and 3D spatial consistency, wouldn't it itself be close to AGI? This presents a "chicken-and-egg" paradox.

Moreover, video models easily overfit in fixed simulation environments, achieving excellent benchmark scores. Once they enter the real world, factors like lighting, time, seasons, or even the opening and closing of curtains affect results, causing a cliff-like drop in performance on actual robots. Essentially, they have never lost a benchmark but find it very difficult to deploy on real machines.

Therefore, I hold reservations about WAM. If one believes in this route, rather than directly building WM, it would be better to first solve video generation. However, those working on video generation are waiting for robots to scale up to return realistic data that conforms to physical laws. This creates a deadlock: I'm waiting for you, and you're waiting for me.

42 Hertz: So, do you think a new architecture will eventually emerge, leading robots toward AGI?

Zhou Shunbo: I think it's possible. But there's also the possibility that the current architectures are already sufficient, because people aren't fundamentally bottlenecked by architecture at this stage.

42 Hertz: Does that mean the bottleneck is in data?

Zhou Shunbo: Exactly. No architecture has truly been fully fed high-quality data yet. Of course, I believe Embodied AI will likely ultimately be an agent system, given its extreme complexity. This system will include foundation models, but how exactly do these foundation models reason? How is their memory managed? How is context handled? All of these aspects may require a complete agent system to maximize the value of foundation models.

42 Hertz: How do you handle data collection?

Zhou Shunbo: Currently, our overall approach focuses primarily on human-collected data; a more accurate term would be 'human-centric'.

Egocentric specifically refers to first-person perspective. Suppose my hands are empty and I'm not using UMI, only wearing an egocentric camera; then the ego plus bare-hand data collected falls under egocentric data. If I use two grippers, such as UMI, then what's collected is ego plus UMI. However, both types of data can be collectively referred to as human-centric data.

Based on our current small-scale experiments, we believe that combining human data with synthetic data is likely the ultimate solution for data.

The first judgment here is that scaling up UMI itself would actually be quite difficult; however, expanding pure ego data faces basically no major obstacles. Ego plus UMI data has already largely replaced real-machine teleoperation data in our current experiments. After all, that was precisely the original purpose behind designing UMI, right?

These data are collected by humans, and the actuators used in UMI are isomorphic to those installed on the robot. Therefore, as long as the spatial trajectory estimation is correct and the workspace boundaries are well-designed, these data can be replayed.

At present, people may still need to wear a dedicated camera on their heads or chests due to considerations such as battery life. However, AI glasses themselves represent an industry trend; ultimately, this data will likely be captured by AI glasses, becoming data naturally generated in daily life.

Therefore, human data combining ego-centric views with bare hands will definitely exist for a long time and become increasingly comprehensive, eventually enabling widespread application in model pre-training.

42nd Radio Wave: There is another view in the industry that Embodied AI is currently stuck on evaluation.

Zhou Shunbo: Yes, I believe the next bottleneck for the industry may be real-robot evaluation. Currently, paths have been found to scale up human-collected data, making it no longer dependent on evaluation.

But next-generation data will likely come from rollout. If evaluation is not done well, closed-loop systems cannot be established, and rollout data cannot be effectively generated. So, after the data problem is alleviated, the bottleneck will quickly shift to evaluation.

42nd Radio Wave: Then how can evaluation be improved?

Zhou Shunbo: This is difficult; currently, no one has an answer. Simulation, world models, and real-robot evaluation each have their own issues, and real robots cannot guarantee completeness.

I believe the best approach is to first clearly define the product boundaries and then deploy it in users' homes for testing. Our concept of 'raising' a robot essentially means outsourcing part of the evaluation process to users.

My current dream isn't to create AGI, but rather to build a robotic product that people enjoy using and which has room for growth. Therefore, my methodology does not pursue theoretical or academic perfection. User evaluations only need to ensure usability within their own homes; whether the product performs comprehensively when moved to a different household is not critical.

42 Hertz:** Will you develop dexterous hands?**

Zhou Shubo: In my view, dexterous hands should be supplied by vendors to robotics companies. If others can produce them well, developing them in-house adds no incremental value; if everyone struggles with them, it indicates the problem is inherently difficult. Given limited resources, startups should avoid investing everything into this area.

For our first version, we will use off-the-shelf two-finger grippers to quickly integrate data, models, hardware bodies, and user teaching processes. Simultaneously, we are designing three-finger grippers in parallel, though currently only achieving three fingers. We cannot halt all prototype and model iterations waiting for an ideal hand.

42 Hertz:** What is your perspective on tactile sensing?**

Zhou Shubo: Tactile technology is still in its early stages. Robotics body manufacturers are waiting for dexterous hands to mature, while dexterous hand companies are waiting for tactile sensors to converge. Tactile sensing remains upstream of even the upstream components, with many intermediate variables involved.

Vision-based tactile sensing theoretically offers the best performance and can reuse visual backbones, but its drawbacks are equally obvious: high computational consumption, slow processing speed, high cost, and fragility. Other array-style or e-skin data formats differ, requiring additional encoder training. Everyone acknowledges the importance of touch, yet there is no consensus, making companies hesitant to bet heavily on any single route.

So the dexterous hand can continue to iterate on its own drive, structure, and reliability, but robot companies don't need to bear all the uncertainty right now. We need to invest our resources in our incremental value.

IV. From the vast ocean to the largest puddle

Radio 42: Do you have sales targets for yourselves?

Zhou Shunbo: What we are currently raising is essentially financial investment, and financial investors do not impose constraints on these matters. My original goal was to achieve at least a thousand units in sales next year. However, some recent developments have made me feel that I need to consider this matter more cautiously.

For consumer-grade robots, on one hand, the reliability, stability, and supply chain of the hardware itself could become bottlenecks. On the other hand, it's about whether the robot is smart or not. In my understanding, regarding intelligence, as long as users can very easily develop fun robots following our current path, that is sufficient.

So, what I am most concerned about right now is whether the entire supply chain allows us to produce this robot reliably and stably, and whether the robot is smart and fun enough.

Radio 42: So the supply chain will still take another one or two years to refine?

Zhou Shunbo: Yes. Relatively speaking, in this wave of embodied AI startups, we came out relatively late. This has both disadvantages and advantages. The biggest advantage is that we can try to avoid the pitfalls that others have already encountered.

Otherwise, I might genuinely feel that setting a target of 1,000 units for next year is somewhat conservative. If I were to shift the company's R&D focus entirely toward ensuring mass production of 3,000 units next year, the entire rhythm of our R&D would become distorted. But in the end, you might be attempting something that is fundamentally unachievable.

**42 Hertz:**So you’re no longer pessimistic?

**Zhou Shubo:**I am long-term optimistic, but there are still uncertainties in the short term.

**42 Hertz:**So do you believe that even without reaching the AGI moment, the current architecture can still be productized and potentially hold certain commercial value?

**Zhou Shubo:**As I said before, I come from a robotics background, so this is how I think about the commercialization of this matter.

We can imagine an axis, moving from left to right, representing the continuous improvement of robot capabilities. Decades ago, industrial robotic arms achieved a commercial closed loop; companies producing AGVs around 2014 have also basically established their business models now. This shows that every step a robot’s capability takes to the right may generate corresponding commercial value.

Today, advances in AI will certainly continue to push robots further to the right, though I dare not say how far it can go.

The problem with the industry today is that a group of people talking about AGI jump directly to the far right, believing that as long as a model is installed into a physical body, the robot can do anything. However, the models have not yet formed true understanding, and the physical bodies are far from being converged. Without even a stable physical body at present, how can we talk about generality? Coupling all complex problems together naturally makes it impossible to implement in practice.

We choose to move forward step by step. We will push as far as our own models allow; if better models emerge in the future, we will leverage them to go further while solidifying the underlying products and engineering.

Therefore, I believe the business model for robots is certain to succeed. Every step holds value; the only difference lies in how large each step can be.

Radio 42: You handle data, models, hardware bodies, and applications, yet you emphasize restraint. What exactly does Euler万象 do, and what does it not do?

Zhou Shunbo: We have a clear internal understanding: embodied AI for homes requires full-stack iteration, but not general-purpose full-stack iteration.

The full stack includes data, models, hardware bodies, and applications. If every link lacks constraints and tries to do everything, it becomes the "general-purpose full-stack" narrative spoken of by many top companies. Our first constraint lies in application scenarios and product definition: clearly defining what we do and what we do not. Once the product is defined, the hardware body can be simplified; once the hardware body and applications are set, the scope for data and models naturally narrows significantly.

At least one thing is very clear: we cannot define ourselves as a pure foundational model company. Current embodied models are not yet worthy of being called true foundational models compared to language models and VLMs. Without running a closed loop across data, models, hardware bodies, and applications, it is premature to talk about foundational models.

Radio 42: What kind of company do you hope Euler万象 ultimately becomes? Since you previously used Bambu Lab as an analogy, are there specific tech companies you particularly admire?

Zhou Shunbo: Yes, I really like DJI and Bambu Lab. I feel that what we are doing now has a floor level comparable to Bambu Lab and a ceiling level comparable to Apple.

There is a logic here. For certain tasks, we have already made the experience of developing robots approach that of Bambu Lab. Users only need to spend a little time getting involved at the beginning, after which they can go play or even sleep. When they come back a few hours later, they can obtain a deployable model and a ready-to-use policy. This is the true consumer-grade experience.

3D printers used to be very complex and cumbersome; the machines were large, and various problems often occurred in the middle, requiring someone to watch them constantly. We can make features like foolproof operation, zero code, and true consumer-grade usability better and better, while supporting an increasing number of tasks. Therefore, I believe we can cover the same market that Bambu Lab covers.

The only difference is that robots are simply much more fun than 3D printers. Even if there are some shortcomings in the early stages, their entertainment value can fully compensate for it. So I consider Bambu Lab to be the floor. Moving further up, this endeavor has the potential to drive an entire ecosystem.

42 Hertz: How many people are on your team currently?

Zhou Shunbo: 52 people.

42 Hertz: 52 people, which hasn't yet reached the scale of the original team you were part of at Huawei.

Zhou Shunbo: No, but it's quite different. At Huawei, about 60% of that team was R&D, while the remaining 20% or so consisted of roles such as solutions, sales, product delivery, and operations and maintenance. Currently, apart from one HR staff member and one administrative staff member, everyone else is engaged in R&D.

42 Hertz: Do you have competitors in mind?

Zhou Shunbo: Currently, I feel that we do not have clear competitors. The main reason is that no company has a natural advantage in the field of home embodied AI. It is a track that requires everyone to redefine and recreate from scratch, and it is highly comprehensive.

Relatively speaking, our team configuration may be more complete, but other companies each have their own strengths. Therefore, I believe that everyone still needs to jointly develop and solidify this track first before competition can even be discussed. We are still in a very early stage.

Of course, I believe this track is large enough. However, since I am not optimistic about AGI or robots that can do everything, it may not be the super-large track it appears to be on the surface.

My personal feeling is that home embodied AI looks like an ocean when viewed from above. As you gradually descend and actually implement solutions, you will likely find at the end that it consists of one puddle after another.

We need to find the largest puddle and dive straight into it. At that point, if others are doing the same thing, we will have competitors. Everyone must first descend together. This is my perception. But right now, everyone is still at this height, and below them all appears to be a vast ocean.