API reference

Contents

API reference#

Environment#

bigym.loco.make(task_name, config=None, **overrides)[source]#

Create a BiGym env; with no overrides, the official configuration.

config is None or an EnvConfig; keyword overrides are EnvConfig fields applied last. See bigym.loco.env.make().

bigym.loco.make_gym(task_name, config=None, **overrides)[source]#

make() behind the gymnasium API (see bigym.loco.gym_adapter).

Configuration#

class bigym.loco.config.EnvConfig(robot_model='g1_dex1', controller=<factory>, episode_length=None, demo_down_sample_rate=10, success_hold_seconds=1.0, reach_tolerance=None, initialization_profile='upstream', enable_all_floating_dof=True, control_pelvis=True, action_mode='absolute', camera_keys=('head', 'right_wrist', 'left_wrist'), camera_shape=(84, 84), state_keys=('proprioception', 'proprioception_grippers', 'proprioception_floating_base'), event_reward_shaping_enabled=False, event_reward_progress_scale=0.0, event_reward_holding_bonus=0.0, event_reward_lift_bonus=0.0, event_reward_target_bonus=0.0, wholebody=None, frame_stack=1, normalize_low_dim_obs=False, action_representation='absolute', upper_delta_scale_rad=None, event_progress_enabled=False, render_mode='rgb_array')[source]#

Every setting an env is built from; the defaults are the official values.

controller=None builds a floating-base env with no lower-body controller (fast, but not the benchmark substrate). episode_length=None takes the task’s registered budget.

Parameters:
  • robot_model (Literal['g1_dex1'])

  • controller (ControllerConfig | None)

  • episode_length (int | None)

  • demo_down_sample_rate (int)

  • success_hold_seconds (float)

  • reach_tolerance (float | None)

  • initialization_profile (Literal['upstream', 'g1_id_v1'])

  • enable_all_floating_dof (bool)

  • control_pelvis (bool)

  • action_mode (Literal['absolute', 'delta'])

  • camera_keys (tuple[str, ...])

  • camera_shape (tuple[int, int])

  • state_keys (tuple[str, ...])

  • event_reward_shaping_enabled (bool)

  • event_reward_progress_scale (float)

  • event_reward_holding_bonus (float)

  • event_reward_lift_bonus (float)

  • event_reward_target_bonus (float)

  • wholebody (WholeBodyConfig | None)

  • frame_stack (int)

  • normalize_low_dim_obs (bool)

  • action_representation (Literal['absolute', 'upper_delta'])

  • upper_delta_scale_rad (float | None)

  • event_progress_enabled (bool)

  • render_mode (str)

override(overrides=None, /, **kwargs)[source]#

Return a copy with the named fields replaced.

Keys are EnvConfig field names; an unknown key raises ValueError. controller / wholebody accept None, a config instance, or a mapping of its fields merged into the current one.

Parameters:
Return type:

EnvConfig

differences(other)[source]#

{dotted field: (other's value, this value)} for every differing field.

Parameters:

other (EnvConfig)

Return type:

dict[str, tuple[Any, Any]]

classmethod from_metadata(metadata)[source]#

The config an env was built from, read back from its env_config block.

Parameters:

metadata (Mapping[str, Any])

Return type:

EnvConfig

class bigym.loco.config.ControllerConfig(backend='groot_wbc_g1', base_action_mode='lowerbody_cmd', pitch_command=True, reset_warmup_steps=200, deterministic_reset=True, init_stance='keyframe', passive_base_tilt=True, default_height_cmd=0.74, init_pelvis_z=0.74, height_cmd_min=None, height_cmd_max=None, pitch_cmd_min=None, pitch_cmd_max=None, default_pitch_cmd=0.0, cmd_clip=1.0, wz_clip=1.0, use_height_cmd=True, skyhook_kp=0.0, skyhook_damping=0.0, model_path=None)[source]#

The lower-body controller in the loop (GR00T-WBC on the G1).

None command bounds mean “the ranges the backend declares”.

Parameters:
  • backend (str)

  • base_action_mode (Literal['lowerbody_cmd', 'legacy_delta'])

  • pitch_command (bool)

  • reset_warmup_steps (int | None)

  • deterministic_reset (bool)

  • init_stance (Literal['keyframe'])

  • passive_base_tilt (bool)

  • default_height_cmd (float)

  • init_pelvis_z (float)

  • height_cmd_min (float | None)

  • height_cmd_max (float | None)

  • pitch_cmd_min (float | None)

  • pitch_cmd_max (float | None)

  • default_pitch_cmd (float)

  • cmd_clip (float)

  • wz_clip (float)

  • use_height_cmd (bool)

  • skyhook_kp (float)

  • skyhook_damping (float)

  • model_path (str | None)

class bigym.loco.config.WholeBodyConfig(mode='hierarchical_current', preserve_leg_proprio=False)[source]#

Outer action over the leg joints too (research modes, off by default).

Parameters:
  • mode (Literal['hierarchical_current'])

  • preserve_leg_proprio (bool)

bigym.loco.config.resolve_config(task_name, config=None, overrides=None)[source]#

The config make builds task_name from.

config is None (the task’s official config) or an EnvConfig used as the base. overrides go on top. A None episode_length becomes the task’s budget.

Parameters:
Return type:

EnvConfig

bigym.loco.config.conformance_violations(config, official)[source]#

Result-affecting fields where config departs from official.

An empty list means the env is the official one for its task, up to fields that cannot change a result (see affects_results()).

Parameters:
Return type:

list[str]

Task registry#

class bigym.loco.tasks.TaskSpec(env_cls, episode_length, data_derived=False, overrides=<factory>)[source]#

One registered task: its class, budget and official-config differences.

env_cls is a BiGymEnv subclass; the env builds it with the BiGymEnv constructor arguments.

Parameters:
config()[source]#

The task’s official configuration.

Return type:

EnvConfig

bigym.loco.tasks.task_config(name)[source]#

The official configuration of a task.

Parameters:

name (str)

Return type:

EnvConfig

bigym.loco.tasks.budget_provenance(name)[source]#

data_derived when the budget came from demos, else upstream_placeholder.

Parameters:

name (str)

Return type:

str

bigym.loco.tasks.all_task_names()[source]#

Every registered task name, sorted.

Return type:

tuple[str, …]

bigym.loco.register_task(name, spec)[source]#

Register a TaskSpec so make(name) builds it.

A name can be registered once; registering it again (a built-in task included) raises ValueError. make("pkg.module:ATTR") builds a TaskSpec without registering it.

Contract layer#

Typed lower-body command declarations.

A lower-body backend consumes a command every control step. Every shipped backend is velocity-conditioned (VELOCITY): the command is a flat vector of scalar channels such as vx / vy / wz / height / torso_pitch. The kind tag leaves room for other kinds of controller (SE3 end-effector targets, reference-motion trackers) without changing the velocity backends.

  • CommandSpec is data: adapters declare their spec. Its bounds are the commands the backend accepts; GR00T-WBC publishes no training ranges, so its adapter declares benchmark clips.

  • height / torso_pitch: the outer action-space bounds derive from the spec unless the config sets them (LowerBody.resolve_command_bounds; the config wins, so recorded demo action spaces do not move).

  • vx / vy / wz: the env clips them with the symmetric cmd_clip / wz_clip; the spec bounds are not read. Per-channel spec bounds would change the action semantics (a substrate bump).

  • rate: an optional slew limit in units/s, None for a policy trained on step commands. It is declarative: adapters own their slew logic.

class bigym.loco.command.CommandKind(*values)[source]#

What a lower-body command vector means.

VELOCITY = 'velocity'#

Flat scalar channels, e.g. [vx, vy, wz, height, torso_pitch…].

groot_wbc_g1 is this kind.

EE_POSE = 'ee_pose'#

Reserved for SE3 hand/foot targets of whole-body controllers.

MOTION_REF = 'motion_ref'#

Reserved for reference-motion trajectories (SONIC-style trackers).

class bigym.loco.command.CommandField(name, unit, low, high, rate=None)[source]#

One scalar channel of a VELOCITY command.

Parameters:
rate: float | None = None#

Slew limit in unit/s; None = step commands (no slew) by training.

clip(value)[source]#

Clip a value into this field’s range (endpoint order-agnostic).

Parameters:

value (float)

Return type:

float

class bigym.loco.command.CommandSpec(kind, fields)[source]#

A backend’s full command declaration.

Parameters:
property dim: int#

Number of scalar channels in the command vector.

property names: tuple[str, ...]#

Field names in command-vector order.

field(name)[source]#

The field with this name; raises KeyError if it is not declared.

Parameters:

name (str)

Return type:

CommandField

has(name)[source]#

Whether the spec declares a field with this name.

Parameters:

name (str)

Return type:

bool

bigym.loco.command.velocity_spec(*, vx, vy, wz, height=None, height_rate=None, torso_pitch=None, torso_pitch_rate=None)[source]#

Build the common twist(+height)(+torso_pitch) VELOCITY spec.

Parameters:
Return type:

CommandSpec

The lower-body controller contract (thin required core).

LowerBodyController is the entire surface the BiGym env integration relies on. Anything else a concrete adapter exposes is an implementation detail. New backends normally subclass bigym.loco.base.LowerBodyBase (the fat base class that absorbs joint addressing, reset pose application, failure detection and declarative replay state) and provide a policy loader, an obs builder and a policy step — a few hundred lines for a real backend — but any object satisfying this protocol plugs in.

Contract summary:

  • controlled_joints joints whose position targets this backend owns.

  • command_spec typed command declaration (see loco.command).

  • control_dt seconds between step() calls (1/control_hz).

  • output_spec names + bounds of the produced targets.

  • reset() re-anchor to the backend’s init pose, clear state.

  • set_command(...) latch the current command (VELOCITY kind channels).

  • step() run the policy once -> joint position targets, in

    controlled_joints order.

  • is_failed() has the robot fallen / left the recoverable region.

  • get_state()/set_state() every mutable field that shapes future

    targets, for bit-exact mid-episode save/restore (demo replay). MuJoCo qpos/qvel/ctrl/qacc_warmstart are snapshotted by the caller; this covers the controller-side remainder (obs histories, last actions, rate-limiter anchors, gait clocks…).

class bigym.loco.controller.OutputSpec(joint_names, low, high)[source]#

Names and bounds of the joint position targets a backend produces.

Parameters:
property dim: int#

Number of joints in this output spec.

class bigym.loco.controller.LowerBodyController(*args, **kwargs)[source]#

Thin required core every lower-body backend implements.

property controlled_joints: tuple[str, ...]#

Joints whose position targets this backend owns (incl. waist).

property command_spec: CommandSpec#

Typed command declaration and its bounds.

property control_dt: float#

Seconds between step() calls.

property output_spec: OutputSpec#

Names/bounds of the produced targets.

reset()[source]#

Re-apply the anchor pose and clear mutable state.

Return type:

None

set_command(cmd_vx, cmd_vy, cmd_wz, *, height=None, torso_pitch=None)[source]#

Latch the VELOCITY-kind command channels.

Keyword names match the command_spec field names (height / torso_pitch). Channels absent from command_spec are ignored. Non-velocity command kinds (EE_POSE / MOTION_REF) will extend this surface as a pure addition; velocity backends stay untouched.

Parameters:
Return type:

None

step()[source]#

Run the policy once; return targets in controlled_joints order.

Return type:

ndarray

is_failed()[source]#

True when the base has fallen / tipped beyond recovery.

Return type:

bool

get_state()[source]#

Snapshot every mutable field that shapes future targets.

Return type:

dict[str, ndarray]

set_state(state)[source]#

Restore a snapshot produced by get_state().

Parameters:

state (dict[str, ndarray])

Return type:

None

get_base_obs()[source]#

(base_lin_vel, base_ang_vel, projected_gravity) in the base frame.

Return type:

tuple[ndarray, ndarray, ndarray]

get_command()[source]#

Current [vx, vy, wz] command (clipped view).

Return type:

ndarray

get_height_command()[source]#

Current height command in meters.

Return type:

float

get_last_action()[source]#

Raw policy action from the most recent step().

Return type:

ndarray

class bigym.loco.base.LowerBodyBase[source]#

Shared pipeline for lower-body backends.

Subclasses set in __init__: controlled_joints, command_spec, control_dt, controlled_range_low/controlled_range_high (from build_joint_ranges()), velocity_clip/yaw_rate_clip, and the latched command, height_command and last_action.

controlled_joints: tuple[str, ...]#

Joints whose position targets this backend owns.

command_spec: CommandSpec#

The backend’s typed command declaration and its bounds.

control_dt: float#

Seconds between step() calls.

property output_spec: OutputSpec#

Names and bounds of the joint targets produced by step().

is_failed()[source]#

Fallen / tipped-over detection from base height and tilt.

Return type:

bool

set_command(cmd_vx, cmd_vy, cmd_wz, *, height=None, torso_pitch=None)[source]#

Default clip-on-set semantics (groot_wbc family).

Keyword names match the command_spec field names.

Parameters:
Return type:

None

build_joint_addresses(joint_names)[source]#

The qpos and dof addresses of joint_names, in order.

Raises:

ValueError – A joint is missing from the model.

Parameters:

joint_names (tuple[str, ...])

Return type:

tuple[ndarray, ndarray]

build_joint_ranges(joint_names)[source]#

The lower and upper limits of joint_names; unlimited joints get +-inf.

Parameters:

joint_names (tuple[str, ...])

Return type:

tuple[ndarray, ndarray]

find_sensor(sensor_name, *, dim)[source]#

The sensordata address of the dim-wide sensor sensor_name, or None.

Parameters:
Return type:

int | None

apply_pose(qpos_addresses, dof_addresses, targets, joint_names)[source]#

Write a joint pose + matching actuator targets and re-forward.

Parameters:
Return type:

None

get_state()[source]#

Snapshot every attribute named in STATEFUL.

Scalars become 0-d float32 arrays, arrays are float32 copies — matching the recorded snapshot format so demo npz files stay interchangeable.

Return type:

dict[str, ndarray]

set_state(state)[source]#

Restore a snapshot; missing keys fall back to _state_default().

Restoration follows STATEFUL order, an ordering contract adapters may rely on: a _state_default for a derived field (e.g. a slewed height that equals the raw command) must come after the field it derives from.

Parameters:

state (dict[str, ndarray])

Return type:

None

get_base_obs()[source]#

Return (base_lin_vel, base_ang_vel, projected_gravity), base frame.

Return type:

tuple[ndarray, ndarray, ndarray]

Backends#

Lower-body backends: the registry and the built-in groot_wbc_g1.

BACKENDS maps each registered backend name to its BackendBinding; register_backend() adds one. A backend name with a colon ("pkg.module:ATTR") instead names a binding in an importable module, imported on first use.

The vendored policy runtime needs only onnxruntime (a core dependency), imported when a controller is built, so the bigym core stays torch-free.

bigym.loco.adapters.register_backend(name, binding)[source]#

Register binding under name (controller={"backend": name}).

A name can be registered once; registering it again (groot_wbc_g1 included) raises ValueError. controller={"backend": "pkg.module:ATTR"} uses a binding without registering it.

Parameters:
Return type:

None

bigym.loco.adapters.resolve_backend_name(name)[source]#

Validate a backend name: registered, or an importable binding reference.

Parameters:

name (str)

Return type:

str

class bigym.loco.BackendBinding[source]#

A lower-body backend’s hooks into env construction.

The env calls the hooks in this order: robot_cls() while building the task env, then build_controller() and configure_model() on the built env. env is the task’s BiGymEnv; a controller may rely on env.model, env.data, env.robot and env.action_space. config is the env’s ControllerConfig: a backend reads the fields that apply to it and may ignore the rest.

Subclass it and set robot_models; override build_controller(). The controller implements LowerBodyController, most easily by subclassing LowerBodyBase.

robot_models: tuple[str, ...] = ()#

EnvConfig.robot_model names the backend drives.

supports_passive_base_tilt: bool = False#

True when the backend can run with the pelvis roll/pitch left passive (ControllerConfig.passive_base_tilt).

robot_cls(robot_cls, config)[source]#

The robot class to build the task env with.

robot_cls is the robot model’s floating-base class. Return it (the default), or a variant with the joints the controller drives actuated and their PD gains set (G1Dex1.variant).

Parameters:
Return type:

type[Robot]

build_controller(env, config, *, control_dt)[source]#

Build the controller on the freshly built task env.

control_dt is the env’s control step in seconds.

Parameters:
Return type:

LowerBodyController

configure_model(env)[source]#

Adjust the compiled model once the controller is built (default: none).

Parameters:

env (BiGymEnv)

Return type:

None

Demos#

Native demo schema v1 (frozen).

A native demo is one teleoperated episode recorded against the controller-in-the-loop env (bigym.loco.env), stored as one .npz per episode plus one metadata.json per collection directory. This module freezes the contract those files satisfy so training/replay code can validate instead of assuming.

Episode npz keys#

Per-step arrays, first axis = outer control steps (T):

  • rgb_obs (T, num_cameras, 3, H, W) uint8

  • low_dim_obs (T, low_dim) float32 — raw (unstacked) layout, see low_dim_component_slices() of the env for the component map

  • action (T, action_dim) float32 — the NORMALIZED outer action in [-1, 1] (commands + upper-body targets + grippers), exactly what the collector fed env.step(). Decode to physical units via the per-dim action_stats min/max recorded in metadata.json (the collection envelope; env.get_demos() adopts it at load). The RAW per-step values live in the diagnostic raw_outer_action array, not here.

  • reward / discount / demo / is_expert (T, 1) float32

  • event_progress (T, 1) float32 — optional

Per-episode scalars / snapshots:

  • seed (1,) int64 — env reset seed

  • pre_engage_steps (1,) int64 — steps between reset and VR engage

  • engage-state snapshot: init_qpos, init_qvel, init_ctrl, init_qacc_warmstart (and init_act when present) — MuJoCo state at the engage moment

  • lb_state.* — the flattened lower-body state snapshot (env.get_lowerbody_state(): ctrl.* controller fields per LowerBodyController.get_state(), env.* env-side fields). Restoring MuJoCo state + lb_state.* reproduces the episode bit-exactly on the Newton-pinned scenes.

metadata.json#

Written once per collection dir by the collector. Required keys:

  • format: "bigym_replay_npz"

  • pipeline_version: collector pipeline version (string, e.g. "2026-08-26-success-hold-training-view-v1")

  • task: the flat env settings (must include episode_length and demo_down_sample_rate — training MUST match these), plus the collection and training success holds

  • lowerbody_policy: the lower-body controller settings, flat

  • optional env_config: the full EnvConfig the batch was recorded in; EnvConfig.from_metadata reads it, and needs it

  • reset_semantics: init keyframe / pelvis z / warmup steps

  • action_semantics: robot model, backend, base_action_mode, stick map

  • optional action_stats: per-dim min/max for [-1,1] rescaling

bigym.loco.demos.schema.validate_episode(episode)[source]#

Return a list of schema violations (empty = valid).

Parameters:

episode (Mapping[str, ndarray])

Return type:

list[str]

bigym.loco.demos.schema.validate_metadata(metadata)[source]#

Return a list of metadata violations (empty = valid).

Parameters:

metadata (Mapping[str, Any])

Return type:

list[str]

npz + metadata.json IO for native demos (schema: bigym.loco.demos.schema).

bigym.loco.demos.io.load_episode(path)[source]#

Load one episode npz into a plain dict (all arrays materialized).

Parameters:

path (Path)

Return type:

dict[str, ndarray]

bigym.loco.demos.io.save_episode(episode, path, *, validate=True)[source]#

Atomically save one episode npz (write-then-rename).

Parameters:
Return type:

Path

bigym.loco.demos.io.load_metadata(demo_dir)[source]#

Load the collection dir’s metadata.json (None when absent).

Parameters:

demo_dir (Path)

Return type:

dict[str, Any] | None

Cut a raw replay-format demo batch to its training view.

An episode succeeds once the task predicate has held for success_hold_seconds; the recorded episode ends on that frame with the single terminal reward (1) and discount (0). The collector records with a longer hold than training and evaluation use, so the raw batch is cut before it trains: each episode is replayed from its engage snapshot on the task’s official env and ends on the control step where that env latches success. A predicate that held long enough, broke and restarted before the collection hold completed makes that step fall anywhere before the raw end, so no fixed trim finds it. The cut goes to a NEW directory and the source batch is never modified. Usage:

python -m bigym.loco.demos.success_hold --demo-dir <batch>
bigym.loco.demos.success_hold.latch_batch(demo_dir, out_dir=None)[source]#

Replay and cut a raw batch; see latch_steps() and cut_batch().

Returns:

The output directory (default <demo_dir>_hold<env hold>s).

Parameters:
Return type:

Path

Hugging Face Hub access to the BiGym 2.0 demonstration dataset.

The public demonstrations live in one Hugging Face dataset repository with one top-level folder per task, each a LeRobot v3 lossless export:

<repo>/
  move_plate/            data/chunk-000/file-*.parquet + meta/ + metadata.json
  reach_target_single/   ...

Two ways to get them onto a machine:

  • Lazy, per task: task_dir() fetches one task’s folder the first time it is needed. env.get_demos() calls it, so running a task pulls exactly that task’s demonstrations (a few hundred MB to a few GB).

  • Ahead of time: bigym-download --all (or download_all()) mirrors the whole dataset; bigym-download --task move_plate pick_box fetches a subset. --local-dir PATH writes the files into a plain folder instead of the cache, which bigym-view --demo-dir PATH opens.

Both go through huggingface_hub.snapshot_download and share its cache ($HF_HOME / $HF_HUB_CACHE; default ~/.cache/huggingface), so a pre-download and a later lazy load never fetch a file twice, and the usual Hub knobs apply (HF_TOKEN for private/gated repos, HF_HUB_OFFLINE=1 to refuse network access and use the cache only).

The repository defaults to DEFAULT_DATASET_REPO; override it with the BIGYM_DATASET_REPO environment variable (and optionally pin a revision with BIGYM_DATASET_REVISION).

exception bigym.loco.demos.hub.DemosUnavailableError[source]#

The dataset repository has no demonstrations for the requested task.

For a benchmark task this means its demonstrations have not been published yet: the dataset is released task by task, and the message lists what is published and what is still pending.

Read a LeRobot v3 lossless task export back into replay-format episodes.

This is the reader half of bigym.loco.demos.lerobot_export. It reads the on-disk format directly (pyarrow + PNG decode); the lerobot package is never imported, so it works on every Python the core package supports.

Only lossless PNG-mode datasets are accepted: video-mode cameras are lossy and would silently change training inputs.

Alignment: the export (meta/alignment.json, version 2) stores transition-scoped features shifted so LeRobot frame k pairs obs[k] with the action executed FROM it; this reader undoes the shift (replay index 0 comes back from the first_transition sidecar, the repeated final frame is dropped), so the episodes come back exactly as the collector wrote them: row t holds the observation at step t together with the action, reward and discount of the transition that PRODUCED it (row 0 is the reset row with a zero action).

bigym.loco.demos.dataset.load_episodes(task_dir, max_episodes=-1, *, action_representation='absolute', upper_delta_scale_rad=None)[source]#

Reconstruct replay-format episode dicts from one task’s LeRobot export.

Returns (source_name, episode) pairs in dataset order. Each episode maps feature names to [T, ...] arrays: rgb_obs [T, cams, 3, H, W] uint8 (camera order from the collector metadata), low_dim_obs [T, D], action [T, A] (normalized outer action), reward, discount, demo, is_expert, event_progress [T, 1] plus any per-step extras the collector stored.

With the default action_representation="absolute" the arrays are the collector’s bit for bit. "upper_delta" derives a training view in memory; the dataset is never modified.

Parameters:
  • task_dir (Path)

  • max_episodes (int)

  • action_representation (str)

  • upper_delta_scale_rad (float | None)

Return type:

list[tuple[str, dict[str, ndarray]]]

Demo collection#

class bigym.vr.collect.config.CollectConfig(task, out_dir=None, episodes=60, seed=1, collect_success_hold_seconds=3.0, episode_steps=None, episode_seconds=None, max_steps=None, keep_empty_session=False, keep_black_rgb=False, store_event_progress=True, store_fullbody=True, yaw_mode='base', base_vx_scale=0.35, base_vy_scale=0.25, base_wz_scale=0.5, base_z_scale=0.004, height_cmd_max=0.8, pitch_rate=0.8, stick_deadzone=0.15, base_cmd_slew=0.7, recenter_height_offset=0.0, vr_space_mode='follow_head', settle_view='curtain', hud='minimal', resolution='lq', mujoco_gl='glfw', openxr_log_level='error', spectator='none', spectator_hz=15.0, spectator_port=8080, spectator_gpu=None, frame_spike_ms=20.0, operator=None, export_lerobot=False, task_text=None)[source]#

One VR collection session (bigym-collect); each field is a flag.

Parameters:
  • task (str)

  • out_dir (Path | None)

  • episodes (int)

  • seed (int)

  • collect_success_hold_seconds (float)

  • episode_steps (int | None)

  • episode_seconds (float | None)

  • max_steps (int | None)

  • keep_empty_session (bool)

  • keep_black_rgb (bool)

  • store_event_progress (bool)

  • store_fullbody (bool)

  • yaw_mode (Literal['base', 'none'])

  • base_vx_scale (float)

  • base_vy_scale (float)

  • base_wz_scale (float)

  • base_z_scale (float)

  • height_cmd_max (float)

  • pitch_rate (float)

  • stick_deadzone (float)

  • base_cmd_slew (float | None)

  • recenter_height_offset (float)

  • vr_space_mode (Literal['follow_head', 'fixed'])

  • settle_view (Literal['curtain', 'live'])

  • hud (Literal['minimal', 'full', 'off'])

  • resolution (Literal['lq', 'mq', 'hq'])

  • mujoco_gl (Literal['glfw', 'egl', 'osmesa'])

  • openxr_log_level (Literal['trace', 'debug', 'info', 'warn', 'error'])

  • spectator (Literal['none', 'mujoco', 'viser', 'both'])

  • spectator_hz (float)

  • spectator_port (int)

  • spectator_gpu (int | None)

  • frame_spike_ms (float)

  • operator (str | None)

  • export_lerobot (bool)

  • task_text (str | None)

Evaluation#

bigym.loco.eval.runner.evaluate(policy, *, task_name, method, checkpoint='n/a', episodes=100, config=None, overrides=None, env=None, verbose=True)[source]#

Run one evaluation block; return summary (+ leaderboard record).

Pass either config / overrides (forwarded to bigym.loco.make() with task_name; the env is created and closed here) or a ready env (kept open).

Returns {"task", "method", "episodes", "success_rate", "episode_rewards", "substrate", "env_config", "protocol_violations", ["record"]}; record (the protocol-v1 leaderboard entry) only for a full block on an official env.

Parameters:
Return type:

dict[str, Any]

The frozen BiGym 2.0 evaluation protocol.

One definition shared by every leaderboard entry:

  • 100 episodes per (task, checkpoint) evaluation;

  • fixed seed block starting at 620000, reserved for evaluation by convention (episode i uses seed_base + i; training-seed overlap is not checked here);

  • checkpoints every 5k training steps, each evaluated on the same block; the reported number is the MEAN OF THE LAST 5 CHECKPOINTS (80k..100k of a 101k budget) +- standard error. peak_final_selection serves the appendix tables. Selecting the maximum on the same test block introduces selection bias; do not use it as the main-table result or treat correlated checkpoints as independent training seeds. This module records individual checkpoints, not that aggregate; disclose checkpoint/seed/episode aggregation when reporting it;

  • an episode is a success when the task reported success and the robot never fell (is_success()). Evaluation must keep reward shaping off.

bigym.loco.eval.protocol.is_success(env)[source]#

Whether the env’s episode so far counts as a success.

Parameters:

env (Any) – An env from bigym.loco.make() or make_gym, or any wrapper that forwards episode_succeeded and episode_fell.

Returns:

True when the task reported success (env.episode_succeeded()) and the robot never fell (env.episode_fell()).

Return type:

bool

bigym.loco.eval.protocol.eval_seeds(episodes=100, *, seed_base=620000)[source]#

The per-episode reset seeds of one evaluation block.

Parameters:
  • episodes (int)

  • seed_base (int)

Return type:

list[int]

bigym.loco.eval.protocol.peak_final_selection(step_to_score)[source]#

Pick the peak-score and final checkpoints from step->score curves.

Parameters:

step_to_score (Mapping[int, float])

Return type:

dict[str, int]

bigym.loco.eval.protocol.protocol_violations(env)[source]#

Why an env’s results are not official; empty when they are.

Compares the configuration the env was built from (env.config) with its task’s official configuration (env.official_config). Every result-affecting field that differs is listed; fields that only change rendering or the policy-side view of observations and actions are exempt (see bigym.loco.config.affects_results()).

Parameters:

env (Any) – An env from bigym.loco.make() or make_gym, or any wrapper that forwards config and official_config.

Returns:

One "field: official X, env Y" line per departing field.

Return type:

list[str]

bigym.loco.eval.protocol.leaderboard_record(*, task, method, backend, success_rate, episodes, checkpoint, substrate_fingerprint, seeds=None)[source]#

One leaderboard JSON entry (schema v1).

Parameters:
Return type:

dict[str, Any]

Coding-agent benchmark#

The coding-agent benchmark: an LLM agent writes the control policy.

An agent works in a sandbox directory containing a task brief, a policy template and a harness that talks to an environment server over a unix socket. The server (bigym.loco.agent.server) owns the simulator, counts the agent’s interaction budget, keeps a ledger and refuses the hidden evaluation seeds. When the session ends, bigym.loco.agent.evaluate scores the submitted policy.py on those hidden seeds with the same episode loop the agent developed against (bigym.loco.agent.episode), the policy running in an interpreter of its own that holds no simulator (bigym.loco.agent.policy_process).

Public surface:

from bigym.loco.agent import EnvTools, EnvToolsConfig, make_env, Tools
from bigym.loco.agent import load_policy, run_episode

EnvTools(task, EnvToolsConfig(...)) builds the environment a policy runs against, with the settings the command line takes.

Names are resolved lazily, so importing this package pulls in neither the simulator nor the optional bigym[agent] dependencies.

bigym.loco.agent.envtools.make_env(task, pitch=None)[source]#

Build the benchmark environment for a task.

Parameters:
  • task (str) – Task name.

  • pitch (bool | None) – None follows the official configuration (the benchmark layout). True or False overrides the torso-pitch command, which is only useful to replay a submission written against the other layout.

Returns:

The controller-in-the-loop env from bigym.loco.make().

class bigym.loco.agent.episode.Tools(env)[source]#

What a policy may call besides reading observations (no reset/step).

bigym.loco.agent.episode.load_policy(path)[source]#

Import path and instantiate the Policy class it defines.

Parameters:

path (Path) – Path to the policy file.

Returns:

A fresh Policy instance.

Raises:

AttributeError – The file defines no Policy class.

bigym.loco.agent.episode.run_episode(env, policy, seed, max_steps=None)[source]#

Run one episode and return its record.

Parameters:
  • env – A client Env or the evaluator’s in-process LocalEnv.

  • policy – An object with reset(obs, tools) and act(obs, tools).

  • seed (int) – Seed handed to env.reset.

  • max_steps (int | None) – Step cap; defaults to the episode’s own time limit.

Returns:

A per-episode record (seed, success, length, reward, termination, fell, wall_s, final_obs).

Return type:

dict