Each policy controls the whole body and the neck of a Unitree G1, using only the onboard camera and proprioception.
Paper overview
Joga is a full-stack system for humanoid dribbling, shooting, passing, and receiving. A torso-mounted actuated neck gives the robot active vision, and data-driven alignment of perception and dynamics preserves the speed and precision learned in simulation when the policies run on hardware.
Abstract
Exhibiting human-like agility in object interaction is a long-standing goal in humanoid robotics. Soccer offers a demanding case study, requiring tight coordination of perception, locomotion, and foot-ball contact. Reinforcement learning in simulation can directly optimize such task outcomes, but sim-to-real gaps in perception and dynamics undermine transfer to hardware. We introduce Joga (Joint Optimization of Gaze and Action), a full-stack system for humanoid dribbling, shooting, passing, and receiving that perceives the ball with an onboard RGB camera on a custom 2-DoF actuated neck. The torso-mounted neck provides rapid, stable active vision, and the whole-body policy jointly controls gaze and action. Rather than relying on broad domain randomization, Joga aligns the simulator to hardware using small amounts of real data. A perception alignment model, trained on 10 minutes of paired onboard-detection and motion-capture data, reproduces the noisy real ball observations. A whole-body unsupervised actuator network, trained on 10 minutes of proprioceptive data, corrects coupled actuation errors and is used to finetune the policy for sharp turns and high speeds. Finally, distributional skill chaining finetunes passing and shooting from states visited during dribbling, enabling transitions from dribbling to either skill. Using only onboard sensing on outdoor fields, Joga reaches an estimated peak dribbling speed of 2.3 m/s (1.7 m/s sustained), reliably completes sharp turns, and achieves higher success than baselines on both individual skills and skill transitions including dribbling to shooting or passing, and from receiving to dribbling or passing. It also performs single-touch receiving of passes at up to 5.2 m/s. These results show that limited real-world data and targeted modeling can preserve agile object interaction learned entirely in simulation.
Real-world skills
Four agile soccer skills, all 31 actuators, onboard sensing only
The problem
Small errors in sight and motion break precise foot-ball contact
Policies trained in simulation see a cleaner ball and a more ideal body than the real robot has. Broad domain randomization hides these gaps but makes the policy slow and conservative.
How can we preserve the speed, precision, and fluidity learned in simulation when a humanoid plays soccer in the real world?
The system
Joga: close both sides of the sensing-action loop
-
1
Active vision hardware
A custom torso-mounted 2-DoF neck carries an RGB camera for rapid, stable gaze changes during fast locomotion.
-
2
Perception alignment model
Learn the state-dependent bias, noise, and dropout of onboard ball detection from paired detector and motion-capture data.
-
3
Whole-body actuator network
Learn joint-position command corrections from 10 minutes of hardware dribbling, then finetune the dribbling policy in the corrected simulator.
-
4
Distributional skill chaining
Finetune passing and shooting from states visited during dribbling so skills hand off mid-stride.
A neck built for the sim-to-real gap
Two servos mounted directly on the torso, with no external linkages, gears, or bearings, so the neck is fast and needs no neck domain randomization.
Active vision while dribbling toward the goal. Inset: the onboard camera on the neck, with the detected ball (green box) and the ball joystick command (orange arrow).
Perception alignment
Learn what the camera actually sees
The Perception Alignment Model (PAM) learns the bias, noise, and dropout of onboard ball detection from 10 minutes of paired detector and motion-capture data, then replaces perception randomization in training.
Dynamics alignment
A whole-body residual for agile dribbling
The Whole-Body Unsupervised Actuator Network (WB-UAN) corrects the simulator's joint-position targets using the state of the whole body, fit from 10 minutes of hardware dribbling with no torque or motion-capture labels. The dribbling policy is then finetuned in the corrected simulator.
Replaying recorded hardware commands in simulation: WB-UAN rollouts (teal) track the hardware reference (white) more closely than the simulator without a residual (orange).
Skill chaining
Strike from mid-stride, not from a standstill
Passing and shooting are finetuned from robot and ball states collected while dribbling, so control can switch skills mid-stride.
Switching from dribbling to shooting in simulation: a striking policy trained from random initial states (left) versus one finetuned with distributional skill chaining (right).
More on hardware
Receiving, ball skills, and playing with people
Questions & answers
Common questions
What does the name mean?
Joga stands for Joint Optimization of Gaze and Action: each policy controls the neck and the body together. It also nods to “joga bonito,” Portuguese for “play beautifully,” a soccer philosophy of creativity, joy, and flair.
Are the videos sped up? Who controls the robot?
No. All hardware videos play in real time (1×). A human operator joysticks the robot live, much like playing a video game: the joystick commands where the ball should go, and the policy decides how the whole body and neck move to get it there.
- D-pad: skill. ▲ walk · ▶ dribble · ▼ receive.
- Left stick. Walk velocity; while dribbling, the target ball velocity. Also aims passes and shots.
- Right stick. Turn.
- R2 (hold). Sprint.
- □ shoot.
- ✕ pass.
Why a custom torso-mounted neck?
The robot must keep the ball in view from long-range approaches down to close-range foot contact, including while accelerating and turning. A torso-mounted 2-DoF neck driven directly by two servos changes gaze quickly and reduces jitter, and avoiding external linkages, gears, and bearings limits the backlash and compliance that would otherwise need to be randomized in simulation.
Why model perception errors in bounding-box space?
The real pipeline goes image → detector box → pinhole back-projection. Modeling residuals on the box (center for bearing, log-size for range) and back-projecting reproduces that interface exactly, so a residual in box size maps directly to a relative depth error, just as it does on the robot.
How does PAM differ from distance-dependent noise models?
A linear baseline conditions only on range with affine per-axis moments and constant dropout. It captures the mean depth bias but not the motion- and pose-dependent structure, so it overestimates depth noise at close range and stays flat against ego-motion. PAM conditions jointly on ball geometry and robot motion and learns a nonlinear distribution over box errors and visibility.
How is WB-UAN different from Contact-UAN?
Contact-UAN learns joint-local corrections from each actuator's own tracking-error history. WB-UAN keeps Contact-UAN's contact-consistent replay and unsupervised fitting, but a single network sees absolute joint positions and commands across the whole body, so each joint's correction can depend on other joints. This better fits recorded hardware trajectories and yields faster, more controllable dribbling.
What runs on the robot at deployment?
Only the skill policies, the onboard ball detector, and proprioception. PAM and WB-UAN are simulator-side models used during training, and motion capture is used only to collect PAM training labels.