Our team

*Equal contribution · Improbable AI Lab, MIT

Paper overview

Joga is a full-stack system for humanoid dribbling, shooting, passing, and receiving. A torso-mounted actuated neck gives the robot active vision, and data-driven alignment of perception and dynamics preserves the speed and precision learned in simulation when the policies run on hardware.

Abstract

Exhibiting human-like agility in object interaction is a long-standing goal in humanoid robotics. Soccer offers a demanding case study, requiring tight coordination of perception, locomotion, and foot-ball contact. Reinforcement learning in simulation can directly optimize such task outcomes, but sim-to-real gaps in perception and dynamics undermine transfer to hardware. We introduce Joga (Joint Optimization of Gaze and Action), a full-stack system for humanoid dribbling, shooting, passing, and receiving that perceives the ball with an onboard RGB camera on a custom 2-DoF actuated neck. The torso-mounted neck provides rapid, stable active vision, and the whole-body policy jointly controls gaze and action. Rather than relying on broad domain randomization, Joga aligns the simulator to hardware using small amounts of real data. A perception alignment model, trained on 10 minutes of paired onboard-detection and motion-capture data, reproduces the noisy real ball observations. A whole-body unsupervised actuator network, trained on 10 minutes of proprioceptive data, corrects coupled actuation errors and is used to finetune the policy for sharp turns and high speeds. Finally, distributional skill chaining finetunes passing and shooting from states visited during dribbling, enabling transitions from dribbling to either skill. Using only onboard sensing on outdoor fields, Joga reaches an estimated peak dribbling speed of 2.3 m/s (1.7 m/s sustained), reliably completes sharp turns, and achieves higher success than baselines on both individual skills and skill transitions including dribbling to shooting or passing, and from receiving to dribbling or passing. It also performs single-touch receiving of passes at up to 5.2 m/s. These results show that limited real-world data and targeted modeling can preserve agile object interaction learned entirely in simulation.

Real-world skills

Four agile soccer skills, all 31 actuators, onboard sensing only

Each policy controls the whole body and the neck of a Unitree G1, using only the onboard camera and proprioception.

High-speed dribbling: 1.71 ± 0.16 m/s sustained, 2.31 m/s peak
Agile dribbling across fields and surfaces
Delicate first touch when receiving
Dribble-to-shoot transition
Dribble-to-pass transition

The problem

Small errors in sight and motion break precise foot-ball contact

Policies trained in simulation see a cleaner ball and a more ideal body than the real robot has. Broad domain randomization hides these gaps but makes the policy slow and conservative.

Perception gaps cause unexpected foot-ball contacts
Actuator gaps cause stumbling

How can we preserve the speed, precision, and fluidity learned in simulation when a humanoid plays soccer in the real world?

The system

Joga: close both sides of the sensing-action loop

  1. 1

    Active vision hardware

    A custom torso-mounted 2-DoF neck carries an RGB camera for rapid, stable gaze changes during fast locomotion.

  2. 2

    Perception alignment model

    Learn the state-dependent bias, noise, and dropout of onboard ball detection from paired detector and motion-capture data.

  3. 3

    Whole-body actuator network

    Learn joint-position command corrections from 10 minutes of hardware dribbling, then finetune the dribbling policy in the corrected simulator.

  4. 4

    Distributional skill chaining

    Finetune passing and shooting from states visited during dribbling so skills hand off mid-stride.

A neck built for the sim-to-real gap

Two servos mounted directly on the torso, with no external linkages, gears, or bearings, so the neck is fast and needs no neck domain randomization.

CAD model of the torso-mounted 2-DoF yaw-pitch neck with the RGB camera
2-DoF neck

Active vision while dribbling toward the goal. Inset: the onboard camera on the neck, with the detected ball (green box) and the ball joystick command (orange arrow).

Perception alignment

Learn what the camera actually sees

The Perception Alignment Model (PAM) learns the bias, noise, and dropout of onboard ball detection from 10 minutes of paired detector and motion-capture data, then replaces perception randomization in training.

Dynamics alignment

A whole-body residual for agile dribbling

The Whole-Body Unsupervised Actuator Network (WB-UAN) corrects the simulator's joint-position targets using the state of the whole body, fit from 10 minutes of hardware dribbling with no torque or motion-capture labels. The dribbling policy is then finetuned in the corrected simulator.

Replaying recorded hardware commands in simulation: WB-UAN rollouts (teal) track the hardware reference (white) more closely than the simulator without a residual (orange).

Snapshot sequence of the robot dribbling through a sharp 135 degree turn
A sharp 135° turn at 1.2 m/s with the PAM + WB-UAN policy.

Skill chaining

Strike from mid-stride, not from a standstill

Passing and shooting are finetuned from robot and ball states collected while dribbling, so control can switch skills mid-stride.

Switching from dribbling to shooting in simulation: a striking policy trained from random initial states (left) versus one finetuned with distributional skill chaining (right).

More on hardware

Receiving, ball skills, and playing with people

Receive-to-dribble
Receive-to-pass
Close-up dribbling
Receiving a pass
Shifty ball skills
Playing with people
Playing with people
Playing with people
Dribbling on grass
Playing with people

Questions & answers

Common questions

What does the name mean?

Joga stands for Joint Optimization of Gaze and Action: each policy controls the neck and the body together. It also nods to “joga bonito,” Portuguese for “play beautifully,” a soccer philosophy of creativity, joy, and flair.

Are the videos sped up? Who controls the robot?

No. All hardware videos play in real time (1×). A human operator joysticks the robot live, much like playing a video game: the joystick commands where the ball should go, and the policy decides how the whole body and neck move to get it there.

PS5 controller with the buttons used to operate the robot highlighted 1 2 3 4 5 6
Photo: Chris Woodrich, CC BY-SA 4.0, via Wikimedia Commons; annotations added.
  1. D-pad: skill. ▲ walk · ▶ dribble · ▼ receive.
  2. Left stick. Walk velocity; while dribbling, the target ball velocity. Also aims passes and shots.
  3. Right stick. Turn.
  4. R2 (hold). Sprint.
  5. □ shoot.
  6. ✕ pass.
Why a custom torso-mounted neck?

The robot must keep the ball in view from long-range approaches down to close-range foot contact, including while accelerating and turning. A torso-mounted 2-DoF neck driven directly by two servos changes gaze quickly and reduces jitter, and avoiding external linkages, gears, and bearings limits the backlash and compliance that would otherwise need to be randomized in simulation.

Why model perception errors in bounding-box space?

The real pipeline goes image → detector box → pinhole back-projection. Modeling residuals on the box (center for bearing, log-size for range) and back-projecting reproduces that interface exactly, so a residual in box size maps directly to a relative depth error, just as it does on the robot.

How does PAM differ from distance-dependent noise models?

A linear baseline conditions only on range with affine per-axis moments and constant dropout. It captures the mean depth bias but not the motion- and pose-dependent structure, so it overestimates depth noise at close range and stays flat against ego-motion. PAM conditions jointly on ball geometry and robot motion and learns a nonlinear distribution over box errors and visibility.

How is WB-UAN different from Contact-UAN?

Contact-UAN learns joint-local corrections from each actuator's own tracking-error history. WB-UAN keeps Contact-UAN's contact-consistent replay and unsupervised fitting, but a single network sees absolute joint positions and commands across the whole body, so each joint's correction can depend on other joints. This better fits recorded hardware trajectories and yields faster, more controllable dribbling.

What runs on the robot at deployment?

Only the skill policies, the onboard ball detector, and proprioception. PAM and WB-UAN are simulator-side models used during training, and motion capture is used only to collect PAM training labels.

Play