ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting

method overview

Illustration of ReWeight, retrieving human demonstrations and assigning sample weights for VLA post-training.

Abstract

Post-training vision-language-action (VLA) models for specific robots and tasks requires in-domain demonstrations, yet collecting diverse robot data is costly. Egocentric human demonstrations provide a scalable alternative, but directly mixing human and robot data can introduce cross-embodiment discrepancies and degrade policy performance. To address this challenge, we introduce ReWeight, a framework that incorporates human data into VLA post-training through demonstration-level retrieval and sample-level weighting. ReWeight learns a cross-embodiment visuomotor representation that combines visual observations with future actions to measure behavioral similarity between human and robot demonstrations. Based on optimal transport, it retrieves human demonstrations relevant to the target robot data and assigns larger weights to samples with smaller cross-embodiment discrepancies. We evaluate ReWeight using π0.5 across eight simulation tasks and four real-world tasks under both clean and randomized settings. In simulation, ReWeight improves the average success rate of post-trained π0.5 from 39% with only robot data and 44% with randomly mixed human-robot data to 57%. In the physical experimental setting, it achieves an average success rate of 68.8%, outperforming the baselines by 28.8% and 13.8%, respectively. Overall, ReWeight provides an effective paradigm for transforming abundant egocentric human experience into transferable supervision for robot learning.
method overview

Overview of the proposed ReWeight framework.

Alignment Visualization Results

Simulation

Stack Bowls
Stack Blocks
Drawer Place
Pick Place Food

Real-world

Stack Bowls
Place Teddy in the Drawer
Place Fruits Basket
Press Stapler

Simulation Results: RoboTwin 2.0

RoboTwin 2.0 Robot-Only Results
RoboTwin 2.0 ReWeight Results

RoboTwin 2.0 — VLA Post-Training: ReWeight selectively retrieves and weights egocentric human demonstrations for post-training, consistently outperforming both Robot-Only training and Mixed Data (Random) human-robot data mixing across eight simulation tasks. The results show that relevant and fine-grained weighted human experience is more effective than indiscriminately adding human demonstrations.

Real-world Results

Below are videos of Robot-Only, Mixed Data (Random), and ReWeight (Ours) on real-world manipulation tasks. (Videos are sped up by 2x.)

real-world results

Real-world evaluation across four physical robot manipulation tasks. The task-level and average success rates show that ReWeight (Ours), which selectively retrieves and weights human demonstrations, consistently outperforms Robot-Only and Mixed Data (Random).

Stack Bowls
Robot-Only
Mixed Data (Random)
ReWeight (Ours)
Place Teddy in the Drawer
Robot-Only
Mixed Data (Random)
ReWeight (Ours)
Place Fruits Basket
Robot-Only
Mixed Data (Random)
ReWeight (Ours)
Press Stapler
Robot-Only
Mixed Data (Random)
ReWeight (Ours)

Real-world Generalization Experiment Results

We further evaluate robustness under varying illumination and visual distractors, comparing Robot-Only, Mixed Data (Random), and ReWeight (Ours) under each setting.

Real-world generalization under environmental perturbations.
Tasks Methods Lighting Distractors Total
Stack Bowls Robot-Only 2/10 2/10 4/20
Mixed Data (Random) 4/10 2/10 6/20
ReWeight (Ours) 7/10 5/10 12/20
Place Teddy in the Drawer Robot-Only 3/10 5/10 8/20
Mixed Data (Random) 6/10 4/10 10/20
ReWeight (Ours) 5/10 5/10 10/20
Place Fruits Basket Robot-Only 5/10 4/10 9/20
Mixed Data (Random) 5/10 7/10 12/20
ReWeight (Ours) 7/10 9/10 16/20
Press Stapler Robot-Only 2/10 0/10 2/20
Mixed Data (Random) 3/10 2/10 5/20
ReWeight (Ours) 5/10 5/10 10/20
Stack Bowls
Robot-Only
Mixed Data (Random)
ReWeight (Ours)
Place Teddy in the Drawer
Robot-Only
Mixed Data (Random)
ReWeight (Ours)
Place Fruits Basket
Robot-Only
Mixed Data (Random)
ReWeight (Ours)
Press Stapler
Robot-Only
Mixed Data (Random)
ReWeight (Ours)
Stack Bowls
Robot-Only
Mixed Data (Random)
ReWeight (Ours)
Place Teddy in the Drawer
Robot-Only
Mixed Data (Random)
ReWeight (Ours)
Place Fruits Basket
Robot-Only
Mixed Data (Random)
ReWeight (Ours)
Press Stapler
Robot-Only
Mixed Data (Random)
ReWeight (Ours)