Paper · arXiv ↗

I-BFM

Reward-Conditioned Robust Humanoid Interaction
via Unsupervised Reinforcement Learning

Ziqi Han1,* Yitang Li2,* Junhan Sun3,* Fanrong Dong3 Yaojie Shen4 Lei Ye5 Zetong Jing3 Yongqi Zhang4,6 Yiming Zhang4 Xue Wang4 Hao Zhao2,4,†
1Tongji University 2Tsinghua University 3Zhejiang University 4RoboParty Lab 5Harbin Institute of Technology 6ShanghaiTech University

* Ziqi Han, Yitang Li, and Junhan Sun contributed equally.   † Corresponding author.

We thank the authors of OmniContact for their help.

Overview

I-BFM / OVERVIEW
I-BFM pre-training and reward inference system overview

Abstract

Behavioral foundation models (BFMs) have recently shown that a single humanoid policy can support diverse whole-body control, but extending such generality to physical interaction remains challenging. We introduce I-BFM, to our knowledge the first BFM for humanoid–object interaction. Rather than relying on task-specific policies or reference tracking, I-BFM learns a shared representation of the coupled dynamics among the humanoid, objects, and their contacts using forward-backward representations and unsupervised reinforcement learning. Given a downstream task reward, the same policy can be directly conditioned on a latent command to execute closed-loop interaction without task-specific policy optimization. To improve interaction control over different time scales, we further train the policy with both short-horizon interaction targets and longer-horizon goal targets. A single I-BFM policy performs carrying, pushing, and kicking, while also supporting goal reaching, motion tracking, stylistic control, and long-horizon task chaining. More importantly, it remains effective after large deviations from nominal execution: on Carry, I-BFM achieves 94.3% nominal success and retains 89.3% success after robot falls, compared with 1.3% for a planning-based baseline. Real-world experiments on a Unitree G1 further demonstrate diverse loco-manipulation behaviors, rapid recovery from interaction failures and external disturbances, and task chaining without task-specific retraining.

Live MuJoCo Session

Approach, lift, transport, and place.

↗

Start Carry Box

Keyboard Control

↗

Start Keyboard Control

Robust Carry Box

Robust Push Box

Robust Kick Box

Robust Goal Reaching

Robustness across
three conditions.

No external disturbance.

Carry Box

94.33%

Push Box

90.00%

Kick Box

81.67%

Box disturbance.

Carry Box

91.00%

Push Box

85.33%

Kick Box

79.67%

Robot disturbance.

Carry Box

89.33%

Push Box

84.67%

Kick Box

86.67%

Reward in.
Behavior out.

Paper · arXiv ↗ Code · Coming soon