JiJia Vision Secures Another 1 Billion RMB in Funding: How World Models Are Entering Homes and Factories?

On June 15, Jiaji Vision announced the completion of a 1 billion RMB Series B2 financing round.
This round of financing was jointly invested by Singaporean cross-border investment institution Lion City Capital (which has continuously followed on in multiple rounds), the China-Belgium Fund, CICC Investment, Wanxiang Qianchao, Fosun Ruiheng, Huagai Chuangying, Jinchuangtou, Deyi Capital, Huacang Capital, Yuanshi Fund, and other national team funds, industrial capital, financial institutions, and state-owned asset platforms. Multiple existing shareholders, including Guozhong Capital, Dacheng Wealth Management, and Turing Asset Management, continued to over-invest.
The proceeds from the financing will be mainly used for continuous investment in the 'double pyramid' data and algorithm system, the research and iteration of physical AGI foundational models, and the scaled implementation in C-end home scenarios and B-end industrial scenarios.
Notably, earlier this year in March and April, Jiaji Vision had already completed a 1 billion RMB Pre-B round and a 1.5 billion RMB Series B1 round, respectively.
In just three months, Jiaji Vision completed a total of three financing rounds, with a cumulative amount reaching 3.5 billion RMB. Such a pace and density of financing are uncommon in the embodied AI industry.
Given that the industry as a whole is still in the early validation stage, why was Jiaji Vision able to complete large-scale financings consecutively this year?
What is being bet on behind this is not just the robot hardware itself, but rather a full-stack route for world models to enter real-world scenarios.

Beyond Actions, Robots Need to Understand the Physical World
In recent years, the embodied AI industry has undergone a noticeable shift.
Initially, the focus was on how robots could move stably; later, it shifted to whether they could complete tasks autonomously.
But as various VLA models began to be applied, an increasingly important question emerged: Do robots truly understand their environment?
The real world is not like a laboratory. Home environments contain cluttered items, changing lighting conditions, and randomly appearing people, full of uncertainties.

Therefore, robots need not only to execute actions but also to form an understanding, prediction, and decision-making regarding the physical world before acting—capabilities that are lacking in language-centric models.
In simple terms, a world model is not a single module but rather a methodology that enables machines to comprehend environmental changes. It resembles the ability of robots to build an understanding of their environment—rather than merely memorizing actions, they attempt to learn the laws governing the real world, thereby navigating into complex, real-world scenarios.
This is precisely why, over the past one to two years, the technical approach of world models has begun to occupy an increasingly prominent position within the embodied AI industry.
Rather than treating the world model as a standalone module, JiaJi ShiJie (Excellent Vision) aims to construct a comprehensive physical AGI capability system.

What is the "Double Pyramid" System?
Against the backdrop of large language models (LLMs) gradually maturing, there is growing expectation to bring AI into the physical world on a large scale, which has spurred the rise of the embodied AI industry.
However, compared to the rapid development of LLMs, embodied AI has yet to fully establish its Scaling Law. The underlying reason lies in the fact that embodied AI lacks both a scalable, describable data system for physical laws and algorithmic architectures capable of efficiently learning these laws.
A core reason for GPT's rapid evolution is the existence of massive text data on the internet. Decades of accumulated web pages, books, academic papers, and code repositories have formed vast training corpora. Model companies typically only need to address how to filter and clean this data, rarely needing to create data from scratch.
However, robots are different; they operate in the physical world, which is a data desert. To overcome this dilemma, robot policy training must rely on diverse data paradigms, including internet video data, human demonstration data, world model simulators, synthetic simulation data, and even real-robot data. The difficulty of acquiring such data is several times greater than that for Large Language Models (LLMs).
Furthermore, even with abundant physical-world data at the current stage, the Vision-Language-Action (VLA) architecture, which is centered on language, struggles to fully digest these data. This is because it typically converts visual frames and action instructions into text-based tokens before feeding them uniformly to large language models for processing.
Such an architectural design finds it difficult to effectively process three-dimensional spatial information, causal relationships in physics, and continuous, coherent action sequences. Consequently, it has inherent shortcomings in understanding the physical world.
Therefore, JiJia ShiJie proposes a method that organizes both data and algorithms into hierarchical structures, forming a "Dual Pyramid" system through deep coupling between a Data Pyramid and an Algorithm Pyramid.

The Data Pyramid essentially attempts to construct a more comprehensive data system. From bottom to top, its five layers consist of: internet video data, human demonstration data, world model simulators, synthetic simulation data, and real-robot data.

The core logic is straightforward: real-world data is expensive, necessitating more low-cost data; however, low-cost data has limited generalization capabilities, requiring constant refinement through feedback from the real world.
In the process of building its data infrastructure, JiJia ShiJie (Excellent Vision) simultaneously launched the wheel-arm robot本体 Shiguang S1, low-cost real-machine data acquisition hardware Maker M01, low-cost handheld data acquisition hardware U-01, and low-cost Ego data acquisition hardware E-01. It also developed its proprietary Embodied World Model platform, GigaWorld-0, forming a complete full-stack software and hardware system.
The company expects to accumulate 1 million hours of real-machine and non-robotic body data by the end of 2026 to support model training and iteration, thereby achieving better generalization capabilities across different tasks and robotic bodies in real-world scenarios.
With data currently being the biggest bottleneck in the embodied AI field, a comprehensive data system often serves as a key competitive barrier among various companies in the industry.
Additionally, there is the equally important algorithm pyramid.

Many current robot models can already generate decent movements, but good performance in demos does not guarantee stable operation in the real world.
The real challenge lies in whether a robot can still complete tasks when the environment changes—such as moving to a different room, altering lighting conditions, or changing object placement. This ultimately comes down to the robot's ability to understand its environment.
Therefore, Excellent Vision's proposed dual-model system of 'world generation + action' essentially aims to create a closed loop that encompasses environmental understanding, change prediction, action generation, and feedback optimization.
Thus, the three layers of the pyramid, from bottom to top, are world simulation, action alignment, and experience reinforcement.
Among them, in the world-action model, GigaWorld-Policy can increase success rates by approximately 30 percentage points on certain tasks, while achieving a tenfold improvement in both training efficiency and inference speed. Meanwhile, GigaBrain-0.5M* can achieve nearly 100% success rate in high-difficulty, long-horizon tasks.
Additionally, in the world generation model, DriveDreamer is the world's first autonomous driving world model oriented toward the real physical world, enabling large-scale deployment of world models in the physical realm.
Notably, JiaJia Vision announced it will release the GigaBrain-1 model in the third quarter of this year and continue advancing the GigaBrain-2 and GigaBrain-3 models. It is reported that GigaBrain-3 will be trained on 10 million hours of video data and 1 million hours of world-action data, aiming for the "GPT-3 moment" of physical AGI.
From data to models, JiaJia Vision has built a self-evolving, mutually reinforcing capability closed loop, during which deployment in real-world scenarios becomes a critical link.

Running Two Lines Simultaneously: Home and Factory
For the embodied AI industry, landing has always been the most frequently discussed topic, after all, technology must ultimately enter the real world to create value.
However, given that the industry is still in its early stages of development, most Embodied AI companies typically prioritize a single direction: either entering factories first or homes first.
The difficulty levels for these two environments are entirely different. Factory environments are more standardized but demand extremely high efficiency, whereas home environments are the most complex yet offer greater long-term potential.
Ji Jia Shi Jie has chosen to pursue both lines simultaneously, advancing its C-end (consumer) and B-end (business) strategies in parallel.
It is reported that its general-purpose humanoid robot, "Shiguang S1," has already secured China's first order for 100 units in a real home scenario. It will begin large-scale operations in the third quarter of this year, while its next-generation home general-purpose robot, "Shiguang S2," will also be released in the third quarter.

Furthermore, in the B-end scenario, in April this year, Ji Jia Shi Jie partnered with FAW Mold and Alibaba Cloud to bring GigaWorld, GigaBrain, and Maker H01 into the real factory environment of FAW Mold.
Validation was conducted around tasks such as palletizing, cross-area transportation, and dynamic obstacle avoidance, compressing the scene adaptation cycle of traditional automation solutions from several months to just a few weeks.

In June this year, the company announced plans to deploy 1,000 general-purpose robots equipped with Jiaji Vision's world model embodied brain and Maker series in Wuxi in collaboration with Longsheng Technology over the next three years.
Against the backdrop of the industry still largely being in pilot and small-scale PoC stages, a deployment of thousands of units is a noteworthy signal.
Beyond creating value through implementation, more importantly, large-scale deployment will generate vast amounts of real-world operational data from robots in various scenarios. This data, in turn, accelerates upgrades to model capabilities, further driving scale deployment and forming a 'data flywheel' of continuously iterating capabilities.

Physical AI Competition Is Not Just About Single-Point Capabilities
Over the past few years, the entire embodied AI industry has showcased numerous impressive demos.
However, there is growing recognition that compared to dance performances or work demos, the truly challenging part lies in enabling robots to enter the real world more stably.
Thus, investors now place greater emphasis not merely on single-point capabilities but on a full-stack closed loop integrating data, models, hardware, mass production, engineering, and scenario closure.
In a sense, the consecutive financing rounds for JiaJia Vision are betting on this very capability.
Huang Guan, founder and CEO of JiaJia Vision, is an Innovation Leadership Engineering Doctorate graduate from the Department of Automation at Tsinghua University. He previously served as Head of Visual Perception Technology at Horizon Robotics and as Partner and Vice President of Algorithms at Jianzhi Robotics. He also has experience working at top research institutions such as Microsoft Research Asia and Samsung China Research Institute. With years of deep involvement in Physical AI, he possesses expertise spanning technological innovation, industrial implementation, and serial entrepreneurship.
Furthermore, the core team's background reflects their journey through the past decade of Physical AI development, covering Computer Vision (CV), autonomous driving, embodied intelligence, and world models. During the CV era, they led multiple globally influential visual AI competitions and won championships. In the autonomous driving era, their BEVDet series of work consistently ranked first globally on nuScenes. In the era of world models and embodied intelligence, DriveDreamer achieved large-scale industrial implementation of world models.
The core team members hail from top academic institutions and leading tech enterprises, including Tsinghua University, Peking University, the Chinese Academy of Sciences, Carnegie Mellon University (CMU); Horizon Robotics, Alibaba Cloud, Bosch, Baidu, and others.
Additionally, a top world model scientist, who is among the top 2% of scientists worldwide with nearly 20,000 citations, and published over 10 first-author papers in top conferences during their PhD, has also joined the team.
Overall, this is a team with ample practical experience in data, models, hardware, mass production, engineering, and scenarios. Such full-stack capabilities are rare in the industry as embodied intelligence gradually moves toward application.

Conclusion
Embodied AI has never lacked stunning demos, with various dance performances constantly emerging and videos of dexterous two-handed operations flooding social media feeds.
But behind the buzz, people are increasingly facing a reality honestly: getting robots to move is never the hardest part. The real challenge lies in ensuring they can stably complete tasks after entering a different room, encountering varied lighting conditions, or being assigned unfamiliar jobs.
This issue points directly to whether an entire full-stack system can truly understand and adapt to the physical world.
Jiia Vision's answer is its "Double Pyramid" architecture, which uses a hierarchical data system to address where data comes from and how it is utilized, while employing a dual-model architecture of world generation and action to solve whether models can genuinely digest physical laws, allowing both to deeply couple and drive each other.
Building on this, C-end home applications and B-end industrial lines are advancing simultaneously, using data flywheels from real-world scenarios to continuously evolve the entire system.
From this perspective, three months, three rounds of financing, and 3.5 billion yuan highlight the core bet on a complete technical roadmap spanning data, models, scenarios, and mass deployment.
