DexRoam Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations

DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations

Rui Zhou1,2,*, Yibo Yuan4,2,*, Junkai Zhao2,*,†, Fangyuan Zhao3, Xiaoguang Zhao5, Shanghang Zhang3,✉, Sirui Han1,✉

1The Hong Kong University of Science and Technology 2Beijing Academy of Artificial Intelligence 3School of Computer Science, Peking University 4Beihang University 5Institute of Automation, Chinese Academy of Sciences * Equal contribution † Project Leader ✉ Corresponding authors

Paper accepted to CoRL 2026

Research Highlight

DexRoam unlocks motion-level human data for mobile bimanual dexterous manipulation by transforming egocentric whole-body demonstrations into a continuous, coupled, robot-compatible action space. It provides a complete pipeline spanning tracker-free egocentric human motion capture, whole-body human-to-robot alignment, temporal action transformation, and VLA-based policy learning, enabling scalable and data-efficient robot learning from diverse human demonstrations while reducing the reliance on costly robot teleoperation data.

Across five real-world tasks and two VLA backbones, aligned human demonstrations raise average success from 29% to 56% on GR00T N1.7 and from 32% to 57% on π0.5, while matching robot-only training with 50% fewer robot demonstrations.

Portable capture

Tracker-Free Egocentric
Whole-Body Human Data Capture System

Portable egocentric capture of RGB and whole-body human motion, without external cameras or body-worn trackers.

Tracker-Free
No external cameras, environment markers, or motion trackers.
Portable Setup
Only a consumer VR headset and a head-mounted stereo camera.
View-Consistent
Live in-headset video streaming aligns the demonstrator’s view with the policy observation.

Robot teleoperation

Whole-Body Mobile
Dexterous Teleoperation

A single operator simultaneously controls locomotion, torso, dual arms, head, and dexterous hands for synchronized whole-body demonstrations.

Whole-body mobile dexterous teleoperation setup connecting the operator’s VR headset, controller, and glove to a mobile bimanual robot
Single-Operator Control
One operator controls the complete robot body.
Coupled Whole-Body Motion
Locomotion and manipulation are demonstrated simultaneously.
Gesture-Based Recording
Start, stop, save, and discard episodes through hand gestures.

Human motion, robot-ready

Human-to-Robot Alignment

Three explicit stages bridge different bodies, control semantics, and execution speeds — without losing the coordination of whole-body motion.

Human-to-robot alignment pipeline: root-centric human motion is mapped to robot-space trajectories and temporally resampled into a robot-executable whole-body action sequence

Embodiment
Alignment

Retarget whole-body human motion into the robot embodiment.

Action-Semantic
Alignment

Express motion as robot-centric relative actions with shared control semantics.

Temporal
Alignment

Resample trajectories by task progress to match robot execution timescales.

Alignment results

Distribution Alignment Visualization

Human and robot whole-body action distributions before and after full alignment.

Pick Chips Can human and robot action distributions before alignment
Before Alignment
Pick Chips Can human and robot action distributions after full alignment
After Alignment

Real-world deployment

Real-world mobile
manipulation deployment

Real-robot evaluation across diverse mobile manipulation tasks.

Pick Chips Can

03 · Policy learning

Learning from
Human Demonstrations

Does aligned human supervision improve policy learning, and which training paradigm leverages it most effectively?

GR00T N1.7 · Task Success Rate

Robot Only Human-Robot Cotrain Human Pretrain + Robot FT Human Pretrain + Cotrain FT

04 · Data efficiency

Reducing Robot
Data Requirements

How does aligned human data improve robot learning with limited robot demonstrations?

50% fewer robot demonstrations. With 25 robot demonstrations, human-assisted training approaches or matches the 50-demo Robot-only reference across the evaluated tasks.

BibTeX

@article{zhou2026dexroam,
  title={DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations},
  author={Zhou, Rui and Yuan, Yibo and Zhao, Junkai and Zhao, Fangyuan and Zhao, Xiaoguang and Zhang, Shanghang and Han, Sirui},
  journal={arXiv preprint arXiv:2609.35761},
  year={2026}
}