Is climbing stairs a huge breakthrough for robots? Skild Brain's end-to-end model evolves machines into humans
Why is a humanoid robot's ability to climb stairs considered a huge breakthrough?
Currently, many humanoid robots still cannot climb stairs smoothly. This fact seems inconsistent with common perception, as many humanoid robots can already dance or perform acrobatics—tasks that are actually difficult for humans to master. So why has the ability to climb stairs become such a significant milestone?
This is a classic example of Moravec's paradox: tasks that are easy for humans are often difficult for robots, and vice versa.
So what makes stair climbing so challenging for humanoid robots?
The most critical factor is the need for precise coordination between visual perception and motor control, requiring dynamic adaptation to variations in step height and geometric shapes.
In contrast, dancing performances typically take place in open spaces and usually do not require visual input; they rely primarily on proprioception and internal motion sensing.
Skild Brain: A Vision-Based End-to-End Locomotion Model
For a long time, there have been two main approaches to humanoid robot walking.
One approach involves creating a map: first modeling the ground as an undulating terrain map, then selecting where to place each step on the map, and finally controlling the legs to follow that path.
The other is the more common locomotion strategy, which is essentially blind walking (not reliant on vision): the robot walks using only proprioception, such as joint angles, leg speed, and whether it has touched the ground. However, this path often leads to problems: once obstacles are encountered, the robot is prone to falling. This is why we often see robots stumbling when climbing stairs.
Recently, however, Skild AI's release of Skild Brain (a vision-based end-to-end motion model) offers a new path. It more closely resembles human walking behavior: walking while looking and adapting, relying on the robot's visual feedback.
This is a locomotion strategy with super strong adaptability, trained end-to-end by a neural network that takes visual perception as input and results in robot motor execution.
In its latest release details regarding Skild Brain, Skild AI primarily demonstrated its low-level control capabilities, which enable fully online vision- and proprioception-driven end-to-end motion control.
Using camera images, Skild Brain can dynamically react to the environment around the robot, with every movement being an instantaneous decision. This allows the model to instinctively adapt to new terrains based on the latest observational information. This similarity to humans lies in how people adjust their movement strategies in real-time when facing changes in different terrain environments.
Strictly speaking, however, Skild AI's locomotion is not purely visual; it utilizes proprioception when vision is obstructed. This is similar to how humans walk: we don't constantly stare at the ground, nor do we fall simply because our view is momentarily blocked.
To test Skild Brain's ability to adapt to environments, Skild AI conducted a test evaluating how the robot autonomously plans paths through obstacles.
Staff constructed an obstacle course featuring unstable carts, unevenly placed wooden planks, and steps of varying heights. The robot had no prior identification of these obstacles, nor were its movements pre-designed. This heavily tested Skild Brain's real-time decision-making capabilities.
The experimental results showed the robot perfectly navigating the obstacles. It adjusted its strategies for different barriers, including decisions on foot placement timing and stride intervals. The entire visual end-to-end system achieved human-like movement instincts.
Subsequently, Skild Brain demonstrated its stair-climbing abilities without requiring any pre-set 'stair mode.' It can adjust its gait according to terrain changes, eliminating the need for specific modes in particular environments—a trait shared with humans who do not rely on complex switching mechanisms.
When navigating stairs, each step is only 3 cm wider than the robot's foot, yet Skild Brain still makes timely and correct reactions to ensure the feet land in the right positions. Throughout this process, the robot shows no hesitation before lifting its feet, and its overall traversal speed remains unchanged—this is the brilliance of Skild Brain.
Even under load, the robot demonstrates stable performance when going up and down stairs.
In terms of reliability in real-world deployment, tests have shown that Skild Brain can operate continuously and correctly during long-range back-and-forth stair climbing without causing the robot to trip.
Additionally, Skild AI tested Skild Brain’s ability to adjust when subjected to external forces. When pushed or pulled while on stairs, the robot quickly adjusts its foothold and maintains balance.
Skild AI: Traditional Generative AI Training Methods Don't Work
Skild AI's technical approach aims to build a continuously improving, general-purpose robot brain applicable across different scenarios, capable of controlling any hardware to perform any task.
To achieve this, Skild Brain employs a hierarchical architecture: high-frequency low-level action policies receive input from low-frequency high-level action strategies, making it suitable for various quadruped robots, humanoid robots, desktop robotic arms, mobile manipulators, and more.
When Skild Brain was launched in July, Skild AI claimed it could drive almost all types of robots, from assembly line robotic arms to humanoid robots.
In terms of its technical path, Skild AI highlights a common yet somewhat overlooked critical issue in the industry. They argue that many so-called robot foundation models lack actionable physical commonsense; in the real world, they often fail to handle complexity.
Many research teams claim to have built so-called robot foundation models by starting with existing Vision-Language Models (VLMs) and incorporating less than 1% of real robot data.
While large language models are rich in semantic information, they are merely superficially polished and lack true underlying operational understanding. Consequently, many systems currently labeled as robot foundation models cannot cope with complex operations in real environments, despite possessing some degree of semantic generalization capability.
In contrast, Skild AI first completes pre-training in simulation environments and human operation videos, then fine-tunes using real operational data from every connected robot. This approach allows them to provide customers with directly deployable solutions.
Skild AI has so far completed three rounds of funding. In June this year, Skild AI just closed its third round of financing, which included $100 million from SoftBank, $25 million from NVIDIA, and $10 million from Samsung, totaling $230 million. This further increased the company's valuation to approximately $4.5 billion.
The core team of Skild AI was co-founded by former Carnegie Mellon University professors Deepak Pathak and Abhinav Gupta, who have been deeply engaged in the fields of robotics and artificial intelligence for over 25 years.
The competitiveness and technical path of the core team are precisely the important reasons why it has been collectively invested in by giants such as SoftBank, NVIDIA, and Samsung. Founded in 2023, this two-year-old team is already a leader in the Embodied AI industry.
Regarding technical details, Skild AI will continue to release information in the coming weeks, which will provide new references for technological development in the industry. Building a universal brain for robots so that they can truly enter real life and cope with complex and changing physical environments obviously still requires much more work.
Reference: https://www.skild.ai/blogs/one-policy-all-scenarios
