
*Equal contribution
TLDR. How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? In this work, we present Do as I Do, an algorithm that reconstructs and retargets monocular RGB human videos to multi-fingered dexterous robotic hands, yielding robot-complete manipulation data from disparate human videos.
How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. However, difficulties in estimating hand-object interaction and crossing the human-to-robot embodiment gap have hindered the adoption of abundant monocular RGB-only human videos as the primary source of robot manipulation data. In this work, we present Do as I Do, an algorithm to reconstruct and retarget monocular RGB human videos to multi-fingered dexterous robotic hands. Do as I Do reconstructs hand-object interactions from various egocentric and exocentric in-the-wild video sources. The algorithm then retargets these hand-object interaction estimates into a sequence of actions executable in the real world, yielding robot-complete manipulation data from disparate human videos. Overall, Do as I Do outperforms previous state of the art in estimating hand-object interactions and extracting dexterous manipulation trajectories from RGB videos, as we show on experiments on datasets with ground truths and on a dataset of video clips collected online. Our experiments enable us to propose an efficacy playbook for practitioners collecting human data for manipulation.
Reconstruction. Visualizations of our hand-object reconstructions, overlaid on the original videos.
Retargeting. Visualizations of our retargeted hand-object interactions, physically simulated in MuJoCo.
We thank Kyutai for providing us with the compute resources for this project. We are grateful to Chaoyi Pan for guidance and insightful discussions on retargeting. We also thank Jane Wu and Hongsuk Choi for helpful advice and discussions on hand-object reconstruction.
@article{paliwal2026doasido,
title={Do as I Do: Dexterous Manipulation Data from Everyday Human Videos},
author={Bhawna Paliwal and Haritheja Etukuru and William Liang and Pieter Abbeel and Nur Muhammad Mahi Shafiullah and Jitendra Malik},
journal={arXiv preprint arXiv:2606.19333},
year={2026}
}