The coding-agent benchmark#
A coding agent gets a one-sentence task, a video of a human demonstration and
a budget of environment steps, and writes policy.py. bigym-agent scores
that file on the same evaluation seeds as the learned policies, through the
same frozen lower-body controller.
Rules and scoring#
A cell is one task × one session. Inside it:
The task statement is one sentence naming what to do and to which object. It says nothing about positions, sizes or the success tolerance.
The demonstration is one human demo rendered from the robot’s own cameras (
head,left_wrist,right_wrist) at 84×84 and 25 fps. There is no outside camera, for the agent or for the policy.The budget is 101,000 environment steps. A reset costs 200 steps (the controller warm-up) plus the steps of the episode. The server refuses the step that would exceed the budget.
The development seeds are the seeds the demonstrations were collected on (60 for a published task), listed in
seeds.json. The server refuses any other seed.The score is the success rate over the 100 hidden seeds of the evaluation protocol. Nothing the agent reports about its own performance is used.
A fall is a failure.
After the 100 evaluation episodes, the first 10 are rerun (
--spot-check N) and must give identical per-episode success. A policy that reads the wall clock or unseeded randomness fails this check.The policy runs in its own interpreter and exchanges observations and actions with the simulator over a pipe. A policy whose code imports
bigym,mujoco,dm_controlorminkis rejected and scores 0.The policy is hand-written control with no learned component: no network, regression or lookup table fitted to data. Tuning constants by running episodes is allowed.
The policy may use numpy, scipy, OpenCV, pillow and the standard library. Nothing else is installed and nothing can be installed.
The agent’s full prompt is
bigym/loco/agent/templates/PROMPT.md.
Each session fills in the task sentence, the seed count, the budget and the
demonstration line. The sandbox also holds docs/api.md (the action, the
observation, the client and the tools the interface grants), the client under
harness/, run_episodes.py, seeds.json, the demonstration videos and a
policy.py template.
Interfaces#
--interface strict is the default and the benchmark row. The policy gets the
learned baselines’ inputs and outputs: the three 84×84 cameras and the raw
low_dim_obs vector in, the physical action out. It has no named state, wrist
positions, camera calibration or IK.
--interface tools is an ablation, reported as a separate row. It adds named
state, wrist positions in the world frame, camera_info, pixel_to_ray and
an IK solver. --no-ik and --no-hand-pos withhold one tool each.
Install#
uv sync --extra agent
export MUJOCO_GL=egl
The agent extra adds OpenCV and scipy, which a submitted policy may import,
and mink for the IK of the tools interface. You also need ffmpeg on
PATH for the videos, and Docker for the container isolation the benchmark
rows use. The commands below run as uv run bigym-agent ....
Run a session#
The agent container can reach only its model API. bigym-agent run creates
the Docker networks and the allowlist proxy and builds the harness image the
first time each is missing.
docker/proxy/README.md
covers the setup, the manual commands and how to pin a CLI version.
Credentials#
Use an API key, in the environment or in a mode-600 file under
~/.bigym-agent/ ($BIGYM_AGENT_HOME):
harness |
API key |
|---|---|
|
|
|
|
Without a key, codex uses the ChatGPT login in
~/.bigym-agent/codex_home (CODEX_HOME=~/.bigym-agent/codex_home codex login)
and claude uses a token from claude setup-token, in
$CLAUDE_CODE_OAUTH_TOKEN or ~/.bigym-agent/claude_oauth_token.
run.json records which mode a session used. A copied host
.credentials.json fails in a container, because host and container refresh
the same token and revoke each other mid-session.
Run#
uv run bigym-agent run --task move_plate --harness codex --model <model-id> --effort high
This writes the cell bigym-agent-runs/move_plate/ in the current directory
(--root puts it elsewhere). A cell holds hundreds of megabytes, so add
bigym-agent-runs/ to your .gitignore.
Each harness runs only its own vendor’s models (codex for OpenAI, claude
for Anthropic), and run refuses a mismatch before anything starts.
--dry-run prints every command the run would execute and writes nothing.
--task move_plate pick_box --parallel 2 runs two sessions at once.
Each session takes the GPU with the most free memory and refuses to start if
it has less than 6000 MiB free (--min-free-mib). --gpu picks a GPU by its
nvidia-smi index.
Watch#
uv run bigym-agent status # live sessions: budget, evaluation progress
uv run bigym-agent status --all # finished ones too
uv run bigym-agent kill move_plate # stop a session, its container and its server
uv run bigym-view --demo-dir bigym-agent-runs/move_plate --follow
With --follow the viewer rescans the cell every few seconds, so new
development episodes and policy versions appear while the agent works.
After the run#
A session whose policy.py is still the untouched template is void and is
not scored. Otherwise run scores the last policy version on the hidden
seeds. An interrupted session is scored and marked interrupted.
--no-eval skips the evaluation, --eval-detached runs it in the background,
and --eval-episodes N shortens it for a smoke test. Reported numbers use
100 episodes.
If the evaluation dies, bigym-agent status shows the cell as
EVAL FAILED (unscored). Score it by hand:
uv run bigym-agent evaluate bigym-agent-runs/move_plate --version <n>
A cell is scored only once eval/vNNN/summary.json exists. evaluate will
not overwrite a finished evaluation without --force.
replay rolls a version out on seeds you choose and writes a viewable batch
and mp4s to replays/vNNN/. It runs only the seeds the directory does not
already hold. --no-video skips the mp4s.
uv run bigym-agent replay bigym-agent-runs/move_plate --version 1 --seeds 620000-620004
summary.json holds the success rate, the interface and image cap the
evaluation ran with, and the substrate fingerprint. protocol_violations
lists the config fields that depart from the task’s official configuration,
and is empty for an official run. A policy that imported the simulator is
named in rejected.
Reading results#
bigym-agent-runs/move_plate/
run.json configuration, verdict, budget spent, tokens and cost
transcript.md what the agent said and ran
policies/ every version of policy.py the agent wrote
eval/vNNN/summary.json the hidden-seed score of version NNN
dev/ the agent's own episodes, per policy version
sandbox/ what the agent saw and wrote
The server snapshots sandbox/policy.py each time its content changes. The
last version is the submission. policies lists every version with its
training and evaluation scores:
uv run bigym-agent policies bigym-agent-runs/move_plate
Training scores come from the server ledger. Each episode counts toward the version that was current when it ran.
The viewer opens a cell, a runs root, or several roots at once:
uv run bigym-view --demo-dir bigym-agent-runs/move_plate
Comparing policies#
--compare plays up to six policies of one task side by side on one seed.
Each slot is <session>:<version>, where the version is vNNN, NNN, last
or submission and the session is matched by name or by a substring that
names exactly one session.
uv run bigym-view --demo-dir runs_s1/move_plate runs_s2/move_plate \
--compare "s1:v008,s2:last" --seed 620003
Adding --record compare.mp4 --record-seconds 10 writes the comparison to an
mp4 and exits. The recording renders in a headless Chrome. --browser none
waits for you to open the page in your own browser, which is much faster on a
machine with a GPU. bigym-view --help lists the remaining flags.
Bring your own agent#
uv run bigym-agent run --task move_plate \
--harness custom --command 'my-harness --prompt PROMPT.md'
The command runs in the sandbox directory, where PROMPT.md holds the task,
with AGENT_SANDBOX, BIGYM_AGENT_PROMPT and BIGYM_AGENT_CELL in its
environment. The server, budget, version snapshots and evaluation are the
same as for the shipped harnesses. Its stdout becomes the cell’s transcript.
A custom command runs on the host under --isolation soft, so nothing
enforces the network allowlist. Check afterwards which tool calls named
anything outside the sandbox:
uv run bigym-agent audit bigym-agent-runs
To drive the pieces yourself:
uv run bigym-agent sandbox --task move_plate --cell runs/move_plate
uv run bigym-agent serve --task move_plate --sandbox runs/move_plate/sandbox \
--ledger-dir runs/move_plate --allowed-seeds runs/move_plate/allowed_seeds.json
# Wait for one "ready" line per worker, run your agent in runs/move_plate/sandbox,
# then send SIGTERM to the server. It records the final policy as the submission.
uv run bigym-agent evaluate runs/move_plate --version <n>
Flags that change a result#
Report these with any result. sandbox, serve, evaluate, record and
run take them under the same names, and a cell’s evaluation reads them back
from its sandbox_config.json.
--interface strict|tools, described above.--image-cap WxH, the policy-facing resolution ceiling. The default is84x84, the resolution of the learned baselines. The demonstration is rendered at the same cap.--slew, a rate limiter on the arm and base commands. It is off by default and not part of the benchmark.
--tier privileged (object positions in the observation) and
--allow-external-cameras (outside views) are debugging aids. A number
produced with either is not a benchmark result.