DynamicHOI: Coupled Dynamics for
Physics-aware HOI Reconstruction

Michigan State University

DynamicHOI refines hand–object trajectories reconstructed from monocular video with physics: hand and object dynamics, coupled through contact forces.

Input estimate (WiLoR / FoundationPose) → DynamicHOI, from the paper. ACCEL in mm/frame².

Abstract

We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanically inconsistent trajectories. Existing methods mainly enforce visual and geometric agreement, leaving the underlying interaction dynamics insufficiently constrained. We propose DynamicHOI, a physics-aware HOI reconstruction framework combining geometry-grounded diffusion refinement with coupled hand-object dynamics. Geometry spatially grounds visual evidence for trajectory refinement, while articulated inverse dynamics and Newton-Euler dynamics derive hand generalized forces and object wrenches for dynamics-level supervision. We further couple hand and object dynamics through contact-force transfer and recover active hand actuation as an interaction-level physical quantity. We formulate its empirical magnitude distribution into a probabilistic prior that penalizes unlikely actuation and suppresses mechanically implausible reconstructed motion.

Experiments on three HOI datasets show consistent improvements in both hand and object reconstruction. The reconstructed trajectories further benefit downstream applications including hand world-model generation and robotic manipulation learning, demonstrating the value of physics-aware HOI modeling beyond reconstruction.

Method

Geometry-grounded Trajectory Refinement

Framework overview of DynamicHOI.
Frozen framewise estimators (WiLoR for the hand, FoundationPose for the object) provide an initial trajectory, which a diffusion-based network refines progressively. Hand and object geometry is encoded and used to spatially ground visual features, and an interaction module models hand–object dependencies before decoding the HOI mesh trajectory.

Coupled Hand–Object Dynamics

Coupled hand-object dynamics.
Articulated inverse dynamics gives the hand generalized forces, and Newton–Euler dynamics gives the object wrench. Contact-force transfer couples the two, so the active hand actuation at accounts for both the hand's own motion and the force it exerts on the object. Its standardized distribution is heavy-tailed (bottom right) and is better captured by a Student-t prior than by a Gaussian, whose negative log-likelihood penalizes unlikely actuation. All physical quantities are computed only during training, so inference adds no extra cost.

Quantitative Results

Comparison on HOT3D, HO3D-v2, and DexYCB (paper Table 1). Bold / underline denote the best / second-best reconstruction methods. † denotes our reproduction; unmarked values are paper-reported. “–” denotes metrics that are inapplicable to the method or unavailable.

(a) Hand–object reconstruction on HOT3D.
HandObject
MethodW↓WA↓PA↓ACCEL↓RTE↓ADD@.1↑ADD@.3↑ADD-S@.1↑ADD-S@.3↑0.1d↑ACCEL↓5°5cm↑
Input estimates (y)
WiLoR†11.743.706.5216.2275.16–––––––
FoundationPose†–––––12.838.925.352.86.570.156.7
Reconstruction methods
HaWoR†11.733.928.3812.9073.69–––––––
WHOLE10.413.266.67–––51.1–69.9–––
DynamicHOI5.261.834.151.6873.4925.865.448.278.47.15.579.8
(b) Hand reconstruction on HO3D-v2.
MethodPA↓AUC↑PA-V↓AUC-V↑F@5↑F@15↑ACCEL↓
Input estimates (y)
WiLoR†7.670.8477.660.8460.6470.9844.12
Reconstruction methods
AMVUR8.300.8358.200.8360.6080.965–
HaMeR7.700.8467.900.8410.6350.980–
Hamba7.500.8507.700.8460.6480.982–
DynamicHOI7.450.8517.490.8500.6560.9852.77
(c) Hand reconstruction on DexYCB.
MethodMPJPE↓PA↓AUC↑ACCEL↓
Input estimates (y)
WiLoR†11.525.330.8936.88
Reconstruction methods
Deformer13.645.22–6.77
HaWoR†11.095.320.8943.88
PAD-Hand10.564.63–3.34
DynamicHOI9.514.680.9063.27

Metrics. HOT3D hand: W / WA-MPJPE (cm; trajectories aligned on the first two frames / the whole sequence), PA-MPJPE (mm), ACCEL (mm/frame²), RTE (%). HO3D-v2: PA-MPJPE of joints / vertices (mm) with AUC in [0, 1], F-scores at 5 / 15 mm, ACCEL. DexYCB: wrist-relative MPJPE and PA-MPJPE (mm), AUC, ACCEL. Objects: ADD(-S) AUC (%) at 0.1 / 0.3 m, 0.1d recall (%), ACCEL (mm/frame²) and 5°5cm success (%).

Hand–Object Reconstruction on HOT3D

Held-out egocentric HOT3D clips. Both hands and the manipulated object are shown from a fixed viewpoint behind the camera wearer. Each estimate is aligned to the ground truth by one rigid transform per clip, fitted on the hands; thin tubes show the recent path of each hand and the object.

Ground truth WiLoR + FoundationPose (input estimate) DynamicHOI (Ours)
Lifting a birdhouse
Lifting a plate with both hands
Picking up a coffee pot
Picking up a toy dinosaur
Carrying a bowl of fruit
Lifting a keyboard with both hands
Pouring with a vase
Scooping fruit with a spatula

Hand Reconstruction on DexYCB

DexYCB test clips (s0 split), played at 0.5× speed. The benchmark scores the hand pose relative to the wrist, so every hand is drawn at the ground-truth wrist position with the ground-truth shape, and is coloured by its per-vertex error (blue = accurate, red = ≥24 mm off). WiLoR's fingers are more often amber or red than those of DynamicHOI. Objects are the ground-truth meshes, drawn in grey for context.

Ground truth Predicted hands: per-vertex error0≥24 mm
Picking up a banana (right hand)hands reconstructed · grey objects: ground-truth meshes
Lifting a foam brick (right hand)hands reconstructed · grey objects: ground-truth meshes
Lifting a wood block (left hand)hands reconstructed · grey objects: ground-truth meshes
Lifting a pitcher (right hand)hands reconstructed · grey objects: ground-truth meshes
Lifting a sugar box (right hand)hands reconstructed · grey objects: ground-truth meshes
Picking up a meat can (right hand)hands reconstructed · grey objects: ground-truth meshes
Picking up a soup can (left hand)hands reconstructed · grey objects: ground-truth meshes
Picking up a power drill (right hand)hands reconstructed · grey objects: ground-truth meshes

Downstream Applications

Two downstream tasks use reconstructed hand trajectories. In each, only the source of the trajectories changes; every other component stays fixed.

Dexterous Manipulation: Human Video → Robot

Task. A UR5 arm with an Allegro hand must pick up an object and lift it into a green target, learning only from one human video. Following ViViDex, the reconstructed human hand trajectory is retargeted to the robot and a PPO policy is trained in simulation to imitate it. The two robots differ only in which method reconstructed the human hand.

Success = the object is within 3 cm of the target while held.

WiLoR's reconstruction jitters (see the wrist-acceleration traces), and the policy that imitates it knocks the object over or fails to lift it into the target. The policy trained on the DynamicHOI reconstruction lifts it into the target. Rollouts use the paper's evaluation protocol (randomized initial object pose).

Success rate (%) ↑
GT hand87.4%
WiLoR25.0%
HaWoR24.9%
HaMeR62.4%
DynamicHOI74.9%

Mean success rate over eight held-out DexYCB relocate tasks, 100 episodes each (paper Fig. 4a).

Egocentric World Modeling: Hand Action → Video

Following EgoHOI, hand trajectories reconstructed on HOT3D condition a video world model that generates the egocentric interaction. Each clip shows the real video (left) and generations conditioned on hand actions from WiLoR (middle) and DynamicHOI (right). The green outline marks where the real hand is; the bottom row zooms on it, and the inset shows the hand condition given to the generator.

Lifting a pitcher. With WiLoR actions the generated hand floats above the pitcher, outside the ground-truth outline; with DynamicHOI actions it grasps and lifts the pitcher as in the real video.
mLPIPS ↓
GT action0.194
WiLoR0.284
HaWoR0.307
HaMeR0.305
DynamicHOI0.260
MPJPE (mm) ↓
GT action24.61
WiLoR38.49
HaWoR36.05
HaMeR44.74
DynamicHOI32.33

Generated-hand appearance (mLPIPS, in the ground-truth hand region) and kinematic accuracy (MPJPE) (paper Fig. 4b).