1X AI VP Eric Jang: Dancing Robots Are Cool, But Useless If They Only Dance at Home
Is a robot that only dances useless?
Recently, Eric Jang, Vice President of the Norwegian-founded robotics startup 1X AI, conducted an in-depth interview with the overseas media outlet The Humanoid Hub.
During the interview, Eric Jang praised robots from companies such as Tesla, Unitree, and Zongqing, repeatedly saying "Cool." He also discussed how the robotics industry places insufficient emphasis on world models, arguing that repeating labor tasks in factories has no future.
When discussing Tesla's Optimus dancing, Eric Jang said "Cool" twice, expressing his approval of robots dancing. However, he pointed out that if a robot can only dance, it would be useless in a home environment.
While most other robot companies focus on more industrial fields, 1X chooses to tackle the most difficult home application scenarios for robots. He stated that the home field is 1X's strength and also aligns with 1X's vision of what true general intelligence needs to possess in the future. If one only repeats a single task in a factory, it is impossible to collect diverse data. Obviously, there are more unstructured scenarios in home environments, and the collected data is richer.
He noted that current humanoid robots cannot go up or down stairs, which limits what they can do at home. Additionally, their range of actions is too limited; most laboratories merely have them stand in front of a table performing certain operations. Tasks such as bending over to pick up items and placing them back on the table, or standing on tiptoes to retrieve items from high shelves, are still lacking.
Compared to the AI built by companies like Tesla, Unitree, and Figure, Eric Jang stated that 1X's main advantage lies in its hardware strategy, adopting a comprehensive 'home-first' approach in its AI strategy aimed at solving reasonable household chores. 1X's Redwood is a vision-language-action (VLA) model capable of controlling Neo's entire body, simultaneously managing the lower body, arms, hands, neck, and more. Furthermore, 1X is among the first humanoid robot companies to truly integrate all elements in a full-body manner.
He pointed out lastly that people have not yet taken world models seriously in the field of robotics. 1X has invested heavily in world models, believing they will become increasingly mainstream in the future.
This interview was compiled by 42 Hongbo without altering the original meaning; the following content is for reference only.
Q: You opened an office in Sunnyvale in the Bay Area and are now moving to Palo Alto. What prompted this move? How is the team adapting to the new office?
A: We are moving into a much larger office where hardware, operations, AI, and software teams will all be under one roof. The primary reason for this is that our team size has grown significantly, and we need to act as quickly as possible. Having everyone concentrated in the same building allows us to move rapidly.
Q: While other humanoid robot companies are addressing more structured industrial application problems, why did 1X choose to tackle the most difficult challenge in the humanoid robotics field—home applications—which is highly challenging both in terms of hardware safety and robotic intelligence.
A: The home domain aligns with both 1X's strengths and our vision of what true general intelligence needs to possess in the future. On one hand, we have a hardware strategy built over the past decade, featuring robots equipped with highly compliant and safe tendon-driven actuators. This is a customer-safety-oriented execution strategy. It enables us to create lighter, safer robots that can coexist with humans and operate in environments where accidental contact might occur. This has always been a core component of our DNA regarding how we envision robots working around humans.
On the other hand, we are becoming an AGI company, not just a robotics company. To achieve this, it is necessary to collect highly diverse data across various environments. Such diversity cannot be obtained from repetitive, single-task operations in factories; we want the data distribution to be as broad and the number of tasks to be as large as possible, and home environments provide this kind of unstructured diversity. Although deploying autonomous systems in such environments is extremely challenging and generalization capability is crucial, it is also the ultimate treasure trove for 1X to collect data, making it worth tackling this problem first.
Q: The 1X AI team released Redwood last week. What was your approach to building it? What prompted you to decide on keeping it compact?
A: Redwood is a vision-language-action (VLA) model that can control the entire body of the Neo robot, including the lower body, arms, hands, and neck, all simultaneously. As far as we know, 1X is one of the first humanoid robotics companies to truly integrate all these elements in a whole-body manner.
This is because we are genuinely focused on household chores, such as being able to complete all operations—from picking up laundry to washing it—in a single model. We really emphasize whole-body control, which has demonstrated to us that our hardware pairs very well with it, and our hardware is also very adept at performing compliant actions, such as grabbing a door and pulling it open. Therefore, we hope to truly demonstrate the practicality of combining our hardware with AI through our AI strategy. That is the key point. Redwood is a language-controlled model; when language input is provided to the model, it can predict various different outcomes.
Q: Many robotics systems decouple high-level vision from low-level reflexive control so that reflexes can run at higher frequencies. Why did 1X decide to integrate high-level vision and motion control into the same channel running at 5 Hz?
A: Redwood is currently not an advanced reasoner, nor does it perform high-level planning. This is also what we are actively working to expand for the next version of Redwood. The current release of Redwood primarily focuses on whole-body control and ensures its completeness in task execution. It can perform rapid low-level control and generalize to different instructions; given various commands, it can operate using its whole body. Many tasks in the home require whole-body coordination, such as squatting to pick up clothes, opening a door, and stepping back.
As development continues around planning-based household tasks, the models will become larger to accommodate these tasks. Nevertheless, 1X has invested significant effort in making the Redwood model as compact as possible. We have experimented with larger models, encapsulating intelligence within a small number of parameters. The key lies in careful training through so-called cognitive prediction tasks, which has greatly helped in reducing the model size.
Q: What are cognitive tasks or cognitive data?
A: These are auxiliary prediction tasks designed to reason about the world, thereby grounding representations so that these representations can truly understand the world. For example, where is my hand, or where will my hand be closed in the future? What do I see regarding objects and similar properties?
Q: How do you scale up data annotation for cognition?
A: We are beginning to see similar work being done in other System 1 and System 2 neural networks. For instance, when training large VLA models, you might jointly train them on internet Q&A tasks (such as visual question answering). Gemini Robotics, for example, may perform such operations at a very large scale, pre-training VLA models and then jointly training them alongside actions.
What 1X does is a scaled-down version of this approach: joint training with auxiliary perception and cognition tasks as well as actions. As it scales up, it may resemble the general internet-scale Q&A tasks being conducted by other major labs. However, how to integrate this into the model so that it runs efficiently while possessing all the intelligence required to complete tasks is something we are currently exploring. I do not believe it will take the form of a massive model paired with a small one.
Q: A significant update from 1X is Neo's reinforcement learning locomotion controller, which enables Neo to climb stairs, run, and walk sideways. How does Neo performing actual household tasks and its mobility capabilities differ from training robots like Optimus to dance via reinforcement learning controllers?
A: The motivation behind 1X expanding and building out the reinforcement learning (RL) stack and conducting natural locomotion updates is to enable AI or teleoperation to maximize the use of Neo to navigate spaces within homes. Currently, most home robots cannot climb stairs; Boston Dynamics' Spot robot is an exception, but that is a quadruped robot.
It remains unclear whether any home robot products can climb stairs, which significantly limits the range of tasks robots can perform in homes. Most laboratories treat humanoid robots as dual-arm manipulators, having them stand in front of a table to perform certain operations. However, the true value of humanoid robots lies in their ability to access all spaces. For example, they can bend down to pick up items from under a table and place them back on top, or stand on tiptoes to retrieve items from high shelves.
Humanoid robots were originally developed because the world is designed for humans—a fact well understood by everyone in the field of humanoid robotics—but their software does not necessarily enable robots to do so. Through this RL update, 1X aims to allow the software to fully leverage the hardware's capabilities to the greatest extent possible.
Regarding Optimus dancing, I suspect it may be similar to the recent demonstrations by companies like Unitree and Zhongqing, focusing more on training reinforcement learning controllers to precisely track specific motion capture trajectories. The robot can follow the trajectory well without falling, making the movement look natural. However, the next step is to make it controllable; if it can only dance, it's not very useful in a home environment because you actually need to control it to perform tasks.
1X hopes that through this update, combining the technology for achieving natural movements via motion capture reference training with the controllability of the first-generation reinforcement learning controller (which can be operated via joystick or VR), will result in robots that are both natural and capable of performing practical tasks. This is the core contribution of this update.
Q: As data volume increases, models may grow. Considering various factors, how do you view increasing onboard computing power?
A: Integrating all these robotic functions—including safety layers, reinforcement learning, various perception tasks for processing sensor data, and potentially advanced autonomous planning and audio generation—into a single GPU is undoubtedly a very tricky problem. Squeezing so many functions into one GPU and scheduling them to avoid GPU contention is challenging. We are indeed very excited about future hardware that allows us to run more programs, but we must also ensure compatibility with our existing hardware. Therefore, we have a dedicated team focused on maximizing the efficiency of our current resources.
Q: What unique challenges does 1X's World Model solve that physics-based simulators like Nvidia Isaac Lab cannot?
A: The World Model can simulate humans, where the humans seen within the model behave in some human-specific ways. In a sense, the World Model implicitly captures human intelligence by generating pixel representations of people. This is a functionality difficult to achieve with traditional rigid-body simulators. Moreover, it excels at handling fabrics and deformable objects, such as moving curtains, lowering, or wiping fabrics. This is also a challenging area for rigid-body simulators. Additionally, it understands liquids, as liquids are items frequently encountered in homes, just like any other object that can be simulated in the World Model.
We built the World Model because we believe that to conduct truly rigorous evaluations in unstructured environments like homes, a data-driven simulator is necessary. 1X needs a data engine that enables evaluation in deployed home environments. This means converting logs from real customer deployments into World Model simulations for counterfactual testing.
Q: Is the World Model essentially a video generation model?
A: The core of the world model is a video generation model that utilizes diffusion technology. However, one significant challenge we face when building this world model is enabling it to be controllable through actions. Most video generation models are text-to-video; you input a prompt, and they generate some video.
We found that the text-to-video type conditional layers did not yield much benefit in the process of converting existing pre-trained text-to-video models into our world model. Consequently, we had to retrain the entire model from scratch to focus on actions. Having a high-quality action dataset was crucial for the success of this work.
Q: 1X mentioned in its blog that even a world model aligned with the real world by only 70% would be a useful tool for candidate policies. As policies mature and the gap between policies narrows, how can you ensure that the world model continuously improves alongside autonomy?
A: This is an excellent open question, and we do not currently know the answer. In my view, the world model allows us to evaluate a vast number of things that we cannot manually design simulations for. It provides very broad coverage and adapts well to our deployment scale. However, it remains uncertain whether it can achieve the 99.9% precise alignment required by physics-based simulators, which is a high bar.
But the world model at least provides breadth, while 1X also uses physical simulation as another tool to evaluate a small number of scenarios with high precision. We hope that by combining both approaches, we can achieve both depth and breadth. 1X will continue to scale up the world model to make it increasingly accurate, but it is difficult to predict where it will reach a plateau.
If it has 90% accuracy, it can distinguish between many different policies across various scenarios; but if it is at 70%, it might only be able to determine whether a policy is significantly worse or better than before, without being able to make finer distinctions in between.
Q: To what extent must a world model reach to become a good synthetic data generation tool?
A: The current designed world model can be viewed as receiving the current state and candidate actions, and telling you what will happen in a general physical manner. To generate data, you do not need to input actions; rather, you want to output actions. You might guide it by prompting "what will happen," thereby feeding text back in. There is an interesting convergence between text-to-video models and action-to-video models; if we can find a way to combine them, we could transition between evaluation and data generation.
We are very confident that 1X currently possesses the data required to generate high-quality video replays. Therefore, generating or diffusing actions may not be too difficult, given the vast volume of image data already available. If we can find a way to achieve highly controllable data generation via text, we will have a solution capable of both evaluating and generating data, which excites 1X greatly. This reminds me of Anthropic's work on "Constitutional AI": they do not pre-train on large internet datasets but instead possess a model for generating data, upon which they then train.
Q: Can a sufficiently advanced world model pass the Turing test of physical reality?
A: It is already visible that human movements in world models are realistic. One can imagine that as scale increases, the "manifested" figures within world models will become increasingly intelligent. Not to mention that the robots themselves become smarter within the world model. If one observes only the behavior of figures in the world model, one will notice distinct intelligent behaviors. I personally am very eager to see whether those figures in the world model can ultimately pass the Turing test. A world model is an interactive model that responds to changes in action. Therefore, the working method of the Turing test might involve the world model responding accurately to interventions in robot actions in a manner consistent with psychophysics.
Q: Beyond robotics, does technology like the 1X World Model hold potential in gaming or immersive experiences?
A: Today, we see some startups essentially attempting to build controllable world models for games. I believe there are many questions surrounding how to package these into products, but certainly, the phenomena observed in world models contain intelligence, especially when they include animals, humans, and robots.
Q: You are a staunch supporter of supervised learning, viewing it as an extremely efficient method of data absorption. Now that reinforcement learning performs exceptionally well in manipulation tasks, do you believe reinforcement learning can play a complementary role alongside supervised learning?
A: I completely agree. However, I still believe that supervised learning is an extremely efficient method for leveraging computational and parameter capacity in deep neural networks. This does not mean I disbelieve in reinforcement learning, which improves systems through interaction and feedback. The question then arises: how to implement reinforcement learning at scale on humanoid robots deployed in the real world, which involves significant infrastructure and algorithmic challenges.
For very large models and the field of general artificial intelligence, choosing the algorithm for absorbing information is crucial; this is not an arbitrary choice. How to make reinforcement learning as much as possible a data absorber is of great interest to 1X, namely, scaling up its own intelligence.
1X has collected a vast amount of data on failures in autonomous systems; integrating this data into its tech stack is a major challenge. The Redwood model is already being trained on both successful and failed rollout data, though not strictly via reinforcement learning (e.g., in policy updates). It leverages failure data more to improve its representations, which can be seen as a first step toward reinforcement learning, with further attempts planned for the future.
Q: 1X has been testing the Neo robot in employees' homes. What preliminary insights have you gained, and how have they influenced decisions in building the AI stack?
A: One key lesson learned from deploying robots in households is that reliability is paramount. This is critical for product experience. I believe reliability constraints are limiting our speed in bringing the robot as a product to homes. If your robot isn't super reliable, you must arrange for maintenance personnel nearby whenever it fails. Therefore, as our business scales—like any company in this field—we essentially must solve the reliability issue to address distribution challenges.
Q: You are looking further ahead, aiming to build truly general-purpose fully autonomous humanoid robots. Do you think you have all the elements needed to scale learning, or do you believe fundamental breakthroughs are still required in algorithmic computing?
A: We certainly haven't figured out every detail yet. I would say our macro strategy—deploying large numbers of safe robots in homes, collecting diverse data, and using that data to train a general intelligence capable of well understanding homes, humans, and the physical environment—is a strategy we believe is heading in the right direction. So, we feel we have that part sorted out.
Of course, many details remain, such as which reinforcement learning algorithms should be used to learn from failure? How to compress all intelligence into small-scale computing devices? How to ensure the safety of bipedal robots? What are the trade-offs between reinforcement learning controllers, autonomy controllers, and planners? Clearly, these issues have not been fully resolved by the company, but the general direction is clear: train deep neural networks based on large amounts of diverse data from human activity in homes, then keep the data flywheel spinning continuously.
Q: Tesla, Unitree, Figure, and many other companies are attempting to build their own AI. What is 1X's unique advantage in building foundational AI for humanoid robots?
A: 1X's main advantage lies in its hardware strategy, which enables deployment in highly diverse environments and allows for accidental contacts due to its safety hardware. When more accidental contact is permitted, more interesting trial-and-error data can be collected, leading to stronger intelligence, because learning from failure is crucial.
Moreover, reinforcement learning controllers are highly robust when trained on failure data, making the collection of high-quality datasets crucial—a key factor in 1X's competitive edge. In terms of AI strategy, 1X adopts a comprehensive "home-first" approach aimed at addressing reasonable household chores. The company possesses a data flywheel that begins with partially teleoperated shared autonomy and will eventually expand to full autonomy. 1X remains steadfastly committed to building general-purpose AI along this trajectory.
In the robotics field, world models have not yet been taken seriously enough. I believe there are companies focused on world models and those focused on robotics. However, just as three years ago I felt humanoid robots were not yet adopted or mainstream, despite widespread discussion about world models, they are not currently central to many AI companies' core strategies.
For instance, frontier labs like Google do not treat world models as the core of their AI systems. Most robotic competitors also do not center their AI systems around world models. However, 1X has invested significantly in this area, anticipating that while it is not yet the core of its strategy, world models will become increasingly mainstream in the future.
Regarding Nvidia's released version of world models, Nvidia's latest Cosmos Predict 2 model features action-controllable capabilities—at least one of the released models does. Nvidia's team has begun incorporating robotics data to enhance controllability through actions.
