Docs
Reading the game
Game intelligence is one of three groups, alongside physical intelligence and coding intelligence.
Summarise the opponent's habits from earlier points and exploit them.
In sport science the ability to read a game is called game intelligence or perceptual-cognitive skill: understanding what is happening on the field, anticipating what comes next, and choosing the action that serves you best. Embodied Agent Olympics asks whether a coding agent shows this ability inside one match, against an opponent it has never seen, and whether it turns what it reads into a better program.
We borrow the classification from sport science.
Robot policies learn motor skills in training and keep them fixed in their weights. Our agents face an opponent they have not seen, write the program that plays, and revise it after every point.
What it means in people
| Part | Meaning | Example from sport |
|---|---|---|
| Situation awareness | Keep scanning the whole field, not just the ball: where teammates, opponents and open space are | elite soccer midfielders turn their heads often before receiving the ball |
| Pattern recognition | Recognise the current situation as a familiar pattern | experts recall a real game position after a few seconds of viewing, but not a randomly shuffled one |
| Anticipation | Judge from early cues what the opponent is about to do | a tennis receiver reads the serve direction from the toss and the backswing |
| Reading the opponent | Remember and exploit the opponent's habits and weaknesses; see through feints | keep playing to a weak backhand; spot a feint in fencing |
| Reading teammates and space | Know where a teammate will go and when you outnumber the defence | a pass played before the teammate starts the run |
| Game management | Adapt the plan to the score and the time left | keep possession when ahead, take risks when behind |
| Decisions under time pressure | Pick the best of several options in very little time | a ball handler deciding to drive, pass or shoot |
Endsley’s three levels of situation awareness: perception (what is there) → comprehension (what it means) → projection (what happens next); then comes the decision.
What it means for an agent
| Part | For the agent |
|---|---|
| Situation awareness | Rebuild the state of play from camera images: positions and velocities of players and ball (perception and state estimation) |
| Pattern recognition and comprehension | Turn the state into game meaning: where the space is, who is unmarked, what the score and the clock imply |
| Anticipation | Predict the opponent's next action, not only the ball's physical flight |
| Reading the opponent | Summarise the opponent's habits from earlier points and exploit them (in-context learning, theory of mind) |
| Reading teammates | Model teammates; coordinate through the team channel |
| Game management and decisions | Choose the strategy that fits the situation; trade thinking time against acting time |
Predicting the ball’s flight is physical understanding and putting the ball where you want it is execution; neither counts as reading the game. Reading the game is about behaviour and the situation: strategy and reading the opponent, and teamwork.
Program play and review
- Program play (during play). The program must carry an opponent model and switch plans with the situation.
- Review (between points). The agent looks at the footage and match record, sums up the opponent and edits its program.
In Athlete mode, the platform supplies the motor skill. The agent handles perception, prediction, decisions, strategy and adaptation. Robotics Engineer mode adds building the motor skill itself. Strategy and teamwork are tested in both modes.
Compared with learned robot policies
| Learned policy (RL, imitation, VLA) | Embodied Agent Olympics | |
|---|---|---|
| How the skill is acquired | large-scale training before deployment | no task-specific training; the agent writes its program during the match |
| Where the skill lives | network weights | a readable program, plus the agent's written reviews |
| Adapting within a match | none (or retraining offline) | review after every point and revise the program |
| Opponent | the environment or opponents seen in training | an opponent the agent has not seen: scripted levels or another agent |
| Time scale | millisecond reflexes of the learned controller | during real-time play, the program runs while the world continues; turn-based sports use a chess clock; between-point reviews allow deliberation and program edits |
| Evaluation | success rate on a task | match results plus component diagnostics (score ladder, ground-truth checks, API and competition tests) |
| Explanation | none | the review states a reason, which we can check against the match record |
Current evidence: Preliminary
| Part | Current evidence | How, and with what n |
|---|---|---|
| Situation awareness (perception) | Yes, in one diagnostic. | Table tennis, ground truth: the program logs its estimates and we replay the match. Ball position median error 0.7–0.8 cm, opponent paddle 3–4 mm (n = 3 models x 3 matches with logging). This part overlaps spatial understanding and physical dynamics. |
| Pattern recognition / comprehension | Partly, from reviews. | LLM-judged reviews (Endsley levels): only 3% of 2,525 model-written reviews (n = 2,525 reviews) identify a pattern in the opponent. Judge is lenient: upper bound. |
| Anticipation (of the opponent) | Partly. | Scouting-report diagnostic in table tennis: predictions about the opponent were no better than a constant-probability baseline (Brier score). Contact-point prediction error was 3.6–4.6 cm (n = 3 models x 3 logged matches); this is ball physics. |
| Reading the opponent | Not directly; shows in results and reviews only. | 12% of 2,525 reviews are opponent-directed (passive analysis, n = 2,525 reviews). A separate prompting/scouting diagnostic used 14 valid matches across 5 conditions (n = 192 reviews across all conditions). Adaptation effect not significant (n = 17 seats, win rate 0.59 → 0.66, p = 0.21). Behavioural test with scripted opponents that have plantable habits: planned. |
| Reading teammates and space | Not yet. | Only soccer 2 v 2 has a team format (one agent per robot, team channel). Metrics (passes into space, space controlled) are planned. |
| Game management | Not directly. | Visible only in match results. |
| Decisions under time pressure | Indirectly. | Real-time sports, and the slow-down factor S as a setting; no dedicated metric yet. |
| Learning during the match | Yes, as an analysis. | Each review can edit the program; program snapshots at every review. Per-edit effect on the next points: not distinguishable from zero so far (n = 807 edits). |
Seeing is not winning: Preliminary diagnostic
Checked against the simulator’s ground truth, the models located the ball to within 0.7–0.8 cm (median) and predicted the contact point to within 3.6–4.6 cm. Those errors did not separate returned from missed balls (n = 135 returned and 31 missed balls, p = 0.92), and the model that saw most accurately won 1 of 6 matches (n = 6). Points are lost in execution and placement.
Setting: table tennis against the scripted opponent, L1 seeds 101 and 111, L2 seed 202, slowed x10, cameras only. GPT-6.1 SOL, Claude Opus 5.5 and GPT-6 Astra: 3 matches with logging and 3 without each (n = 18 matches), plus 3 logged matches of the reference player. The diagnostic measures perception and physical prediction.
Sport-science references
- A. M. Williams and K. A. Ericsson (2005). Perceptual-cognitive expertise in sport: some considerations when applying the expert performance approach. Human Movement Science.
- D. T. Y. Mann, A. M. Williams, P. Ward and C. M. Janelle (2007). Perceptual-cognitive expertise in sport: a meta-analysis. Journal of Sport & Exercise Psychology.
- M. R. Endsley (1995). Toward a theory of situation awareness in dynamic systems. Human Factors.
Nine abilities
We name nine abilities in three groups: physical intelligence, game intelligence and coding intelligence.
We mark, for every sport, which abilities it mainly tests.
During play, the 10 real-time sports keep running while the player thinks. The 8 turn-based sports give each player a chess clock (60 minutes per match).
Physical intelligence
Only cameras, no coordinates; physics computed by a real engine; actions carried out by a robot with real joints, torques and tracking lag. During real-time play, the world never waits.
| Ability | Definition | Typical sports | How it shows |
|---|---|---|---|
| 3D spatial understanding | Recover 3D from camera images only: position, depth and orientation, camera geometry, occlusion, fusing several views; reason about angles, lines and slopes in 3D | table tennis (triangulating from two cameras), billiards (bank lines), putting (slopes), foosball (figures hide the ball) | points lost to misjudged positions; score gap in a perception ablation (planned) |
| Physical dynamics | Predict how things move (flight, drag, spin, bounce, friction, collisions, wind) and infer parameters (how fast this table is, how strong the wind is) from observation | table tennis and tennis (spin, bounce), badminton (heavy drag), curling and bowling (friction, collision chains), archery and disc golf (wind) | landing-point prediction error; stability when physics parameters vary (planned) |
| Embodiment and execution | Know what one's own robot can do (reach, speed, torque) and turn intent into accurate action; in the Robotics Engineer mode, also build the control stack (trajectories, inverse kinematics, tracking) | all sports | hit rate and precision; points lost as out of reach |
| Time and latency | Know how old the image is, how long one's own thinking and program take, and how long the arm needs; act for the world as it will be when the action lands | all real-time sports | timing error at contact; points lost as late; score versus slow-down factor |
Game intelligence
| Ability | Definition | How it shows |
|---|---|---|
| Rules | Know how points are scored and what is a foul, and use the rules (walks in baseball, the hammer in curling) | fouls; choices on key points |
| Strategy and reading the opponent | Judge the situation, trade risk against reward, prepare several moves ahead; recognise an opponent's habits and exploit them | more wins later in a match against an opponent with habits; quality of the position after a shot |
| Teamwork | Share roles, pass into space, move to the right place, talk through the team channel | passes to an open teammate; space controlled; team-channel use |
Coding intelligence
| Ability | Definition | How it shows |
|---|---|---|
| Real-time programming | Write perception, prediction and strategy as a program; debug it during the match; fix it quickly when it breaks | program ready at the start; crashes and timeouts |
| Learning during the match | After each point, find why it was lost and change the program or the plan | scoring before and after a review; what was changed and whether it helped |
Programming and learning during the match are tested by every sport.
Coverage
| Ability | Sports with primary focus |
|---|---|
| 3D spatial understanding | 14 |
| Physical dynamics | 14 |
| Embodiment and execution | 18 |
| Time and latency | 8 |
| Rules | 6 |
| Strategy and reading the opponent | 9 |
| Teamwork | 1 |
| Real-time programming | 18 |
| Learning during the match | 18 |
Teamwork is covered by one sport today (soccer 2 v 2); basketball 3x3 is in development.
Marks are declared per sport by the authors from the rules and the play. Filled dot = 2 (primary focus), open dot = 1 (involved), empty = 0 (hardly involved).
Physics tags
Collisions 7 · rolling and friction 7 · air drag 6 · spin and bounce 4 · ballistics 4 · wind 3 · inertia and actuator lag 2 · contact 2 · walking 1.
How a match works
Watch, code, play, revise, grade
During real-time play, the world never waits: thinking costs game time. After the match the recording is replayed and graded in a separate container. In the Engineer mode the agent’s own skill is then tested on hidden requests.
- Watch. The agent sees the game only through the sport’s cameras, with no coordinates. It reads the rule constants and its robot’s specification (
game spec). - Code. It writes perception, prediction, strategy and the commands for its robot.
- Play. Real-time sports never wait for the player: the world runs at a declared slow-down factor and thinking costs game time. Turn-based sports give each player a chess clock (60 minutes per match).
- Revise. After every point the world freezes for a review (
game ready WHY PLAN, at most 5 minutes in real-time sports): the agent reads what happened, edits its program, and says why. - Grade. The recording is replayed in a separate grading container, its state hashes are checked, and the winner is decided there. Players cannot touch the grader.
Every match is recorded and replays bit-exactly, so any match can be checked, analysed or re-rendered later.
Athlete and Robotics Engineer
Athlete. The agent has the basic motor skills of the sport. A fixed control system provided by the platform moves the joints, keeps the balance and carries out the motion. This mode mainly evaluates perception, prediction, planning, decisions and competitive strategy.
Robotics Engineer. The agent receives the robot’s physical model, a low-level control interface and general-purpose basic capabilities, such as an IK library or the vendor’s walking policy. It may implement, combine or improve control methods. This mode mainly evaluates motor-skill development: trajectory planning, controller design, tool reuse, integration and debugging.
| Athlete | Robotics Engineer | |
|---|---|---|
| Platform provides | the sport's skill API (table tennis: hit(t, pos, normal, vel)) | robot model, joint interface, basic tools (IK, walking policy) |
| Agent delivers | a match program | a match program and its own implementation of the same skill API |
| Tested by | matches (against scripted opponents, or duels) | the match itself, then two post-match tests of the final implementation in an isolated container: an API test (hidden requests: timing, position, orientation, velocity errors) and a competition test (our fixed reference strategy plays levels L1–L4 on top of the agent's implementation) |
| The robot | fixed: one body per sport for every agent | a research variable: compare agents on one body, then repeat on others |
Table-tennis control chain: hit decision (when, paddle pose, velocity) → trajectory → inverse kinematics → control and tracking → joint servos and physics. The Athlete owns the hit decision; the Engineer owns the chain through control and tracking. Joint servos and physics are the same for both, with real limits enforced.
Reference lines: the reference player built only from public tools; a boundary baseline that uses the tools but develops no skill (for example, holding the paddle in the ball’s path with IK, which must score about 0); and doing nothing (0).
Where the modes exist today
- Current tasks for table tennis, tennis, badminton and fencing declare both action levels: skill API and joint targets. The other 14 sports have one action level, their own native commands.
- The full Robotics Engineer protocol, including post-match API and competition tests, is a shared library with a minimal example sport. For real sports it has run in pilots: air hockey, fencing, a Unitree G1 shooting drill, table tennis on Unitree G1 and on Rainbow RB-Y1.
- Both modes are defined for every sport and have been validated in pilots.
Air hockey has a separate two-mode pilot with Athlete skills and Engineer joint plans. It uses simulator state as a pilot simplification. The current air-hockey sport uses camera images and native mallet-target commands.
Racket sports get both modes. Soccer gets Athlete only; sprint and hurdles (planned) get Engineer only. Release-device and machine sports get Athlete only unless a core skill remains. A sport gets Engineer mode only if the boundary baseline scores about 0.
Program play and direct play
In program play, the default everywhere, the player writes a program. In direct play, with code execution switched off, the model issues one action per reply. Direct play exists only for table tennis today; direct play with Engineer mode is not offered.
Four formats
A format defines who plays whom, seats per side, side rotation and ranking. Points, sets and fouls belong to the sport’s referee. Each sport declares the formats it supports.
| Format | Seats | How it is played | How it is ranked | Current sports |
|---|---|---|---|---|
| Single | 1 | against a scripted opponent at levels L1–L4 (L1–L2 or L1–L3 in some sports), or with no opponent at all | by the result at each level, or by the score | all 18; the only format of archery, basketball shooting, bowling, disc golf and putting |
| Duel | 2 sides x 1 seat | two players head to head; each pairing plays both sides / both first-player orders, on fixed seeds | wins | 13: the 10 real-time sports, billiards, curling, darts |
| Team | 2 sides x k seats | one player per robot; each sees only its own robot's cameras; teammates talk only through a logged, rate-limited team channel the opponents cannot see (it can be switched off to compare) | team wins | soccer 2 v 2 (basketball 3x3 in development) |
| Multi-party | N sides x 1 seat | all players in one event at the same time | places converted to points | robot racing (4 cars) |
- A coach layout, with one player commanding the whole team, is formally a duel. Soccer’s duel format uses this layout.
- Sports where players do not interact keep only the single format; darts keeps a duel because racing to finish has strategy.
Games rules
DraftDraft. These medal rules are not decided.
Single, duel, team and multi-party formats each have their own ranking. The medal table counts golds, then silvers, then bronzes, never a total score. Reference players and baselines run alongside but take no medals.
- Each Games is a frozen version. An event is one sport, one format, one mode.
- Every event awards gold, silver and bronze. The medal table is ordered by golds, then silvers, then bronzes.
- Athlete events and Engineer events award medals separately, in two medal tables. Body-swap matches are demonstration events without medals.
- Reference players and baselines never take medals; they are shown alongside to mark what is possible and where zero is.
Open decisions
- Knockout rounds for duels or round-robin only.
- Medals for single events against scripted levels.
- Engineer events ranked mainly by the competition test.
Season 1: Exhibition season
Exhibition season · earlier rules · 100 of 300 planned matches · ties share medals.
Four models, 11 sports, 100 valid matches (n = 100 matches).
One framework
Embodied Agent Olympics tests physical intelligence, game intelligence and coding intelligence.
One small core runs every sport. A sport is a folder: its world, its actions, its referee and its observations, declared in one specification file. Players, a model plus its harness, talk to the core through one protocol. The core runs the match, applies the format and rules, records everything and decides who won. Adding a sport touches only its folder; adding a model touches only the players.
| Part | One line |
|---|---|
| Players | a model plus its harness and play setting; knows no sport, cannot reach match internals |
| Sports | a world, its actions, a referee and the observations; never imports another sport or changes the core |
| Core | runs every match the same way: flow, rules, recording, judging; contains no sport or harness name |
| Tools | turn sports and players into tasks, matches and seasons, and recordings into videos and reports; never affect a result |
Three contracts connect them: the player protocol (player ↔ core), the sport interface (core → sport), and the match record (core → tools). The middle is narrow and versioned; both sides grow freely.
One core, 18 sports
- Every sport uses the same core for match flow, clocks, review and judging.
- Single, duel, team and multi-party are core rules: a sport declares which ones it supports.
- A shared robot library uses real robot models and real limits.
- Grading replays the recording in a separate container and checks state hashes. Every match replays bit-exactly.
Harbor
Matches run on Harbor. Single-player tasks run on stock Harbor; duels use a multi-seat extension.
Each seat has its own container, network and account. The multi-seat extension is separate from stock Harbor. Table tennis, badminton and air hockey single-player tasks run on stock Harbor. Grading runs in a separate verifier container.
Bit-exact replay
Physics is deterministic and the core records every command with its time step. Replaying a match reproduces every body pose bit for bit; the grader checks the state hash. This lets us re-render a match, analyse it against ground truth or audit a result.
The physics engine is fixed per sport: MuJoCo, pooltool, or a validated sport-specific model. Renderers are plug-ins, and every match replays bit-exactly, so any match can be re-rendered. Isaac Sim RTX is used for spectators’ offline re-renders.
Real limits and physical feasibility
- Real limits (B0). Every body obeys the real robot’s joint position, velocity and acceleration limits; where the vendor SDK and URDF disagree, the SDK wins; unpublished values are marked as estimates. Platform skills and scripted opponents obey the same limits.
- Physical feasibility (B9). Before a sport-and-body pair counts, an upper-bound player that knows the true state must be able to do the sport’s basic actions under real limits. In racket sports, it must return at least 90% of legal serves. Pairs that fail change the setting or move to a challenge track.
Fencing, air hockey and badminton now enforce arm limits through the shared limits module. Table tennis and tennis have joint-target range and speed checks; their arm controllers do not yet use that module.
Air hockey uses KUKA’s joint ranges, velocities and torques, Air Hockey Challenge acceleration limits, and estimated jerk limits. A KUKA torque-rate limit is not published and is not enforced.
Badminton uses BWF court lengths × 0.4 and a 1.524 m net. Its Panda arm follows the robot’s limits; its x-y gantry is limited to 5 m/s and 50 m/s². The rules declare 9 deviations.
Held to real-robot limits is the rule of the benchmark; this does not mean every sport already passes the audit.
How to add a sport
Adding a sport touches only its folder.
- Write the sport folder: world, actions, referee, observations, built-in players and
sport.toml. - Run the generator and the checks: specification, generic tests, golden replay and fairness of both seats.
- Submit the sport for review.
A sport never imports another sport or changes the core.