1X Unveils Latest Progress on World Model: Can NEO Robot Successfully Enter Homes?

1X is accelerating the large-scale deployment of its NEO robot.

On January 13, embodied AI company 1X officially released 1XWM, a generative world model pre-trained on internet videos. Its core objective is to enable robots to form their own understanding of the world and continuously learn and act upon this foundation, thereby advancing the large-scale deployment of the NEO robot.

After 1X launched the NEO robot last year, the most pressing question was how it would be deployed in home environments. Official introductions noted that for tasks NEO cannot perform, 1X experts could intervene via teleoperation.

This approach has also led many to question the autonomous capabilities of its robots, but today's launch of 1XWM, the world model by 1X, brings a new logic paradigm that offers fresh perspectives for the industry on robotic autonomy.

Let Robots Act First in Their Minds

Traditional VLA (Vision-Language-Action) models are largely reactive, moving from perception input to policy networks and then to action output. While this structure performs well in scenarios with clear rules, it often suffers from instability in complex environments such as homes.

The core operating logic of 1XWM has undergone a paradigm shift: robots no longer react directly to reality but first construct an internally predictable world replica.

For example, after receiving a specific instruction, 1XWM generates text-driven video simulations to envision the entire task process. It then uses an inverse dynamics model to convert these videos into joint commands, enabling the robot to reproduce the imagined behaviors.

Under this logic, when faced with instructions or environmental changes, the robot first perceives the current state of the world through multimodal inputs such as vision and language. It then simulates potential future changes within its internal world model, evaluates the consequences of different action paths in the imagination space, and finally maps the optimal actions back to the real world for execution.

This logic essentially introduces the human cognitive mechanism of 'thinking clearly before acting' into robotic decision-making.

However, it is important to note that 1XWM does not aim to build a precise digital twin of reality. Instead, it focuses on a compressed representation of the world's key causal structures—identifying which changes are stable, which actions lead to irreversible consequences, and how the environment evolves over time.

In summary, the robot first visualizes its successful state and then executes accordingly.

What is the Core of 1XWM?

For a model, suitable hardware serves as the carrier for its full potential, making software-hardware co-design a key focus for many robotics companies. 1X emphasizes the importance of "humanoid" form in its system design, treating hardware as part of the model's distribution.

When robots deviate too much from humans in terms of joints and compliance, the physical priors learned from videos are prone to failure. This gap in capability translation can only be narrowed when hardware is regarded as a core component of the AI technology stack.

Here, 1X also stated that NEO's hardware can basically replicate the process envisioned by the model.

The technical team then introduced that 1XWM's backbone is built upon a 14-billion-parameter generative video model. To adapt this model to NEO's body structure, a multi-stage training strategy was adopted.

First, the model was trained on 900 hours of first-person human videos to align with first-person manipulation tasks. In this stage, the model could capture general manipulation behaviors but struggled to generate videos of NEO performing tasks. Subsequently, it was fine-tuned based on the body structure to adapt to NEO's visual appearance and kinematic features.

By combining pre-training with embodiment fine-tuning, over-reliance on data can be reduced and the robot's generalization ability enhanced.

However, in practical execution, various issues arise. For instance, world models sometimes generate videos that defy physical common sense, such as teleportation or excessive bending.

In this regard, 1XWM first generates highly realistic videos, which are then converted into commands by an IDM (Inverse Dynamics Model). This process filters out unscientific visual artifacts, preventing the robot from executing unreliable actions depicted in the videos.

Ultimate Goal: Entering Households

During actual testing, 1X demonstrated 1XWM's generalization capabilities, handling unseen tasks such as grasping unfamiliar objects and interacting with humans in real-time. After incorporating a Best-of-N strategy, its success rate on delicate tasks like 'drawing a smiley face' was also improved.

However, during operation, generating a 5-second video with 1XWM takes 11 seconds. Such latency could impact specific tasks when deployed in home environments.

Additionally, the technical team found that the quality of videos generated by the world model significantly affects task success rates. To address this, they implemented a strategy of generating multiple videos and selecting the highest-quality one, which increased NEO's success rate in the 'tissue extraction' task from 30% to 45%.

For NEO, a humanoid robot designed for home use, execution speed and success rates are critical during deployment. Therefore, 1X must establish a self-evolving flywheel to ensure continuous improvement of the 1XWM powering NEO.

Last month, 1X formed a strategic partnership with private equity giant EQT. Together, they aim to deploy up to 10,000 NEO humanoid robots manufactured by 1X within EQT’s global portfolio companies between 2026 and 2030.

Deploying 10,000 humanoid robots within five years requires 1X to solve the pressing model-related challenges facing the robotics industry today.

Through the flywheel created by 1XWM, NEO can explore and refine its own strategies without relying on expert demonstrations. This self-improvement mechanism brings NEO closer to its ultimate goal of working in households.

Although only half of 2026 has passed, it is evident that the entire robotics industry is accelerating its efforts toward practical deployment. More technologies are being designed around long-term operation and self-evolution, as systems capable of continuous growth in the real world hold true potential for large-scale application.