Xiaomi's robot model scales like an LLM using 100K hours of gripper video
Robotics has lagged language and vision AI because it lacks the massive, clean datasets that drive scaling laws. Xiaomi-Robotics-1 attacks that bottleneck by borrowing the LLM playbook: a broad pre-training stage followed by targeted post-training. Pre-training runs on 100,000 hours of “embodiment-free” UMI trajectories — handheld-gripper video spanning 1,700-plus household, commercial, industrial, and outdoor scenarios — auto-labeled by a vision-language model that describes how each short clip changes the scene. Post-training then aligns that general action-generation ability to real hardware using cross-embodiment data, including 7,200 hours collected in actual homes, and shifts the model from following scene-transition descriptions to executing plain natural-language instructions.
The headline claim is that robot policy learning obeys clean scaling behavior. Validation action error drops steadily as data and model size grow during pre-training, and — more importantly — that improvement carries through post-training to real-robot success rates in unseen environments with unseen objects, showing no sign of saturating. A stronger pre-trained model reliably yields a better physical robot out of the box, which is the property that has been missing from robotics and is the reason the approach matters.
Downstream, the model adapts to brand-new complex tasks like phone packing, printer refilling, and laundry loading from just a few hours of demonstrations each, hitting a 75% success rate under 10 hours per task — roughly double the π0.5 baseline — and 85% with a larger budget. It also reports state-of-the-art results across four mainstream simulation benchmarks. The work is a proof of concept that cheap, robot-agnostic demonstration video can substitute for scarce real-robot data as the fuel for a general-purpose manipulation foundation model.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.