ByteDance Unveils New VLA Model, Companion Robot Becomes a Household Helper
On July 22, ByteDance's Seed team released a new VLA model, GR-3, which supports high generalization, long-horizon tasks, and dexterous manipulation of flexible objects with dual arms. Also unveiled was the general-purpose dual-arm mobile robot, ByteMini.
What are the highlights of the GR-3 and ByteMini released by the Seed team? Among them, GR-3 possesses the ability to generalize to new objects and environments, understand language instructions containing abstract concepts, and manipulate flexible objects with precision. It can be efficiently fine-tuned with only a small amount of human data, enabling rapid and low-cost migration to new tasks and recognition of new objects. This differs from previous VLA models that required extensive robot trajectory training.
Thanks to an improved model architecture, GR-3 can effectively handle long-horizon tasks and perform highly dexterous operations, including two-handed cooperative manipulation, manipulation of flexible objects, and full-body operations that integrate chassis movement.
These capabilities are achieved through a diverse model training methodology: in addition to high-quality real-world data collected via teleoperated robots, the team also征集ed (collected) human trajectory data based on VR devices with user authorization, and conducted joint training using publicly available large-scale vision-language data—the integration of diverse data is one of the highlights distinguishing GR-3 from existing VLA models.
In these two products, GR-3 is positioned as the "robot brain," while ByteMini is the companion robot designed specifically for it.
As a general-purpose dual-arm mobile robot with high flexibility and reliability, ByteMini serves as the "flexible body" crafted specifically for the "brain" known as GR-3.
The robot features 22 degrees of freedom overall, including an unbiased 7-degree-of-freedom mechanical arm. Close observation reveals that the arm's wrist employs a spherical design, enabling precise operations in confined spaces.
In terms of perception, ByteMini is equipped with multiple cameras: two on the wrists for detailed views and one on the head for global awareness. For motion, it utilizes a Whole-Body Control (WBC) system. As the robotic body, ByteMini integrates the GR-3 model, allowing efficient handling of complex tasks in real-world environments.
GR-3 demonstrates three key characteristics across various tasks: 'mindful,' 'dexterous,' and 'strong generalization.'
In long-sequence table-clearing tasks (with ≥10 subtasks), GR-3 achieves high robustness and success rates while strictly following step-by-step human instructions. When facing multiple identical items (e.g., several cups), it can place them all into a trash bin as commanded. If a command is invalid (e.g., 'put the blue bowl in the basket' when no blue bowl exists on the table), GR-3 accurately recognizes this and remains stationary.
In complex and dexterous clothing-hanging tasks, it can control the coordinated operation of both arms to manipulate deformable flexible objects, robustly recognize and organize clothes placed in different orientations, and stably handle situations where clothes are disordered.
In various object grasping and placement tasks, it can generalize to grasping unseen objects and understand instructions containing complex abstract concepts. For example, during the process of hanging clothes, it can generalize to short-sleeved garments that were not included in the training data.
In terms of technology, GR-3 adopts a MoT network architecture, integrating the "vision-language module" and the "action generation module" into a 4 billion parameter end-to-end model. Regarding data training, GR-3 breaks through the limitations of traditional robots that only learn "robot data," adopting a three-in-one data training method. During training, it can simultaneously acquire knowledge from three data sources: robot data obtained via teleoperation, human VR trajectory data, and publicly available image-text data.
According to reports, ByteDance's Seed team plans to expand the model scale and training data volume in the future, and introduce RL methods to further improve generalization and break through the limitations of existing imitation learning.
As a key indicator for measuring the quality of VLA models, generalization capability enables robots to break boundaries in complex and changing real-world scenarios and quickly adapt to new tasks. As robot companies successively launch VLA models and continuously exert efforts on the 'robot brain', generalization capability is undoubtedly one of the focal points of R&D.
