Daily Average of One Financing Deal: The Crazy July for Embodied AI and the Signals Behind It

In the first half of July, there were 16 financing rounds in China's Embodied AI sector, averaging one per day—a frequency that is staggering.

The total financing amount exceeded RMB 3 billion. In comparison, throughout all of 2024, the Embodied AI field saw a total of 90 financing rounds with a combined amount of RMB 8.933 billion. So where did the money go? 42nd Radio Waves has compiled an analysis.

From these financing data, it can be seen that more than half of the enterprises' financing was concentrated in Series A and Series B stages. Among companies that disclosed their financing amounts, those such as Xingdong Era, Hanyang Technology, Deep Robotics, Xiaoyu Smart Manufacturing, Tashi Zhixing, Xinghai Tu, and Kuawei Intelligent all surpassed the RMB 100 million mark. Investors include giants like Meituan and Didi; looking slightly earlier to June, we also see Tencent and Alibaba. It is no longer surprising for tech giants to cluster as "financial backers" in Embodied AI.

The intensive entry of big tech companies brings multi-dimensional support, including capital and resources, to Embodied AI enterprises.

Embodied AI Concerns the Next Wave of AI

Actually, from the perspective of long-term industry trackers, capital's intensive investment in Embodied AI is not accidental or a case of following trends, but an inevitable trend in industrial development.

Observers who continuously monitor capital flows should have noticed a signal: capital is gradually shifting its focus from large models to Embodied AI. According to data from IT Juzi, there were 72 financing events in the large language model (LLM) sector in the first half of 2025, compared to 44 in the same period of 2024. In contrast, Embodied AI saw 116 financing events, up from 30 in the previous year. Clearly, the growth rate of LLM financing events lags behind that of Embodied AI.

This reflects deeper strategic thinking by capital.

Dr. Xing Dadi of Benzhen Capital stated that the large model track has basically reached its closing stage. Leading companies have established their advantages, and it is unlikely for new players to enter.

The competition in large models is no longer the fragmented landscape of the 'Hundred Models War' in 2023. The stimulating effect of newly released large models is diminishing marginally over time. So far, leading large models have divided the user base and occupied competitive advantages. For new players hoping to rise prominently in this environment, the difficulty is high; they must either achieve technical breakthroughs or have sufficient financial flexibility ('burn money'). However, training large models requires massive amounts of hardware resources such as GPUs and TPUs, which offers poor cost-effectiveness for investors. Embodied AI, on the other hand, is still in its early betting stage, making the shift in capital understandable.

Secondly, we are moving towards the next wave of AI, with Embodied AI being key. In the waves of Generative AI and Reasoning AI, we have witnessed AI's ability to understand, decompose, and generate information. However, in most cases, they remain merely 'tools', not 'life'.

In this context, AGI—which can self-learn, self-improve, and self-adjust to solve any problem without human intervention—is far from arrival. To reach AGI, we must accumulate the ability to interact with the physical world. Therefore, 'bringing artificial intelligence into the physical world and thereby driving the physical world' becomes a more compelling narrative than the aforementioned 'tools', launching the grand narrative of Embodied AI.

When AI truly learns to walk steadily, operate dexterously, and coexist safely in the physical world, it will have taken the crucial step from an 'intelligent tool' to an 'intelligent partner'.

Jensen Huang has publicly stated on this matter: "The next wave of artificial intelligence is Physical AI. AI can enter physical machines, such as robots." Embodied AI serves as a crucial pathway for AI to transition into the physical world within this context.

Meituan, which is heavily investing in embodied AI, shares similar considerations. From Wang Xing's perspective, artificial intelligence must enter the physical world, and Meituan acts as the connector between the digital and physical worlds. Its investments in the embodied AI sector and collaborations with invested companies are key strategic moves to realize this vision.

VLA Has Become Industry Consensus

In this recent wave of funding for embodied AI, we can clearly discern a signal: walking on two legs—"brain" and "body"—has become the mainstream approach, with the "brain" gaining increasing prominence. Among the leading embodied AI companies that have secured the most funding rounds and amounts this year, most have chosen this dual-track strategy prioritizing the "brain," differing from the past where hardware held absolute dominance.

In this strategic direction, the VLA model (Vision-Language-Action model) has become a consensus in the embodied AI industry. It is regarded as a universal architecture connecting perception, language, and behavior, enabling robots to weave language intent, visual perception, and physical actions into a continuous decision-making flow.

Karol Hausman, Co-founder and CEO of Physical Intelligence, stated that VLA is a critical cornerstone for achieving general intelligence, allowing robots to learn from multi-source data such as the internet and translate it into concrete actions.

Having completed seven funding rounds with a cumulative total exceeding 1 billion yuan in less than a year and a half since its inception, X Square Robot has bet on VLA (Vision-Language-Action) models from the start. From initially outputting only actions, it now integrates outputs of actions, language, vision, and chain-of-thought reasoning. The company’s adherence to the "unified end-to-end large model" approach for both brain and cerebellum functions is precisely the core reason why multiple investment institutions favor X Square Robot.

Furthermore, the pace at which domestic robotics companies are deploying VLA models has accelerated this year. For instance, Agibot released its first general-purpose embodied foundation model, Agibot Qiyuan (GO-1), in March. Adopting the ViLLA architecture composed of a VLM and MoE, it enables learning from human videos to achieve rapid few-shot generalization. In June, Galbot launched its self-developed product-level end-to-end navigation large model, TrackVLA, which features pure visual environmental perception, language-command-driven operation, autonomous reasoning capabilities, and zero-shot generalization ability.

Today's VLA is no longer just a simple model but rather a rapidly evolving mindset: enabling robots to directly 'read the world' and 'move,' serving as the key bridge connecting artificial intelligence and the physical world.

Continuous capital injection and deepening involvement in the Embodied AI industry actually represent a process of returning AI to its intelligent origins: intelligence does not exist solely within abstract symbols but should be rooted in continuous interaction between the body and the physical world.

Empowering AI with the ability to perceive environments, understand physics, plan actions, execute tasks, and learn from them in real physical spaces constitutes the ultimate goal of Embodied AI. Enabling AI to better enter the physical world and serve humanity is also the grand vision driving the frenzy of capital flowing into Embodied AI. In the journey toward the next wave of AI, no one will stop.