2026 · Xiaomi Robotics · Advanced
Xiaomi-Robotics-1: Scaling VLA Models with 100K Hours of Real-World Trajectories
Xiaomi-Robotics-1 studies how VLA-style robot policies scale when pre-trained on over 100,000 hours of real-world manipulation trajectories.
Direct answer
What does Xiaomi-Robotics-1: Scaling VLA Models with 100K Hours of Real-World Trajectories contribute?
Xiaomi-Robotics-1 studies how VLA-style robot policies scale when pre-trained on over 100,000 hours of real-world manipulation trajectories.
Background
The system uses large-scale embodiment-free UMI trajectory pre-training, automatic language annotation of state transitions, and real-robot post-training. Xiaomi reports scaling behavior across data volume and model size, plus strong results on RoboCasa365, VLABench, and RoboDojo.
Problem
The work addresses a central constraint in VLA: building systems that learn useful representations or actions while remaining general enough to transfer beyond a single demonstration or environment.
Core idea
Xiaomi-Robotics-1 studies how VLA-style robot policies scale when pre-trained on over 100,000 hours of real-world manipulation trajectories.
Architecture and method
The system uses large-scale embodiment-free UMI trajectory pre-training, automatic language annotation of state transitions, and real-robot post-training. Xiaomi reports scaling behavior across data volume and model size, plus strong results on RoboCasa365, VLABench, and RoboDojo.
- 100K+ hours of UMI real-world manipulation trajectories
- VLM-assisted auto-labeling pipeline
- Real-robot post-training and benchmark gains
Results and impact
It is one of the clearest 2026 examples of the robotics data bottleneck being attacked with human-collected real-world manipulation data instead of relying only on robot teleoperation or simulation.
Prerequisites
- VLA models
- Robot datasets
- Imitation learning
Recommended reading order
Read the explanation above, review the related topic pages, then use the primary-source links below to inspect the abstract, figures, experiments, and released implementation.
Primary sources
External links are provided after the context needed to evaluate the work.
Follow-up research
Related papers and concepts
RT-1: Robotics Transformer for Real-World Control at Scale
RT-1 trains one transformer policy on a large multi-task dataset of real robot demonstrations.
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
RT-2 co-trains vision-language models on web and robot data so semantic knowledge can influence actions.
Open X-Embodiment and RT-X
Open X-Embodiment combines robot datasets across institutions and trains policies that transfer across embodiments.
OpenVLA: An Open-Source Vision-Language-Action Model
OpenVLA is an open 7B-parameter VLA trained on the Open X-Embodiment dataset.
Common questions
Frequently asked questions
What is the main idea of Xiaomi-Robotics-1: Scaling VLA Models with 100K Hours of Real-World Trajectories?
Xiaomi-Robotics-1 studies how VLA-style robot policies scale when pre-trained on over 100,000 hours of real-world manipulation trajectories.
Why is Xiaomi-Robotics-1: Scaling VLA Models with 100K Hours of Real-World Trajectories important?
It is one of the clearest 2026 examples of the robotics data bottleneck being attacked with human-collected real-world manipulation data instead of relying only on robot teleoperation or simulation.
What should I learn before reading Xiaomi-Robotics-1: Scaling VLA Models with 100K Hours of Real-World Trajectories?
Recommended prerequisites are VLA models, Robot datasets, Imitation learning.