API reference#
Environment#
Configuration#
- class bigym.loco.config.EnvConfig(robot_model='g1_dex1', controller=<factory>, episode_length=None, demo_down_sample_rate=10, success_hold_seconds=1.0, reach_tolerance=None, initialization_profile='upstream', enable_all_floating_dof=True, control_pelvis=True, action_mode='absolute', camera_keys=('head', 'right_wrist', 'left_wrist'), camera_shape=(84, 84), state_keys=('proprioception', 'proprioception_grippers', 'proprioception_floating_base'), event_reward_shaping_enabled=False, event_reward_progress_scale=0.0, event_reward_holding_bonus=0.0, event_reward_lift_bonus=0.0, event_reward_target_bonus=0.0, wholebody=None, frame_stack=1, normalize_low_dim_obs=False, action_representation='absolute', upper_delta_scale_rad=None, event_progress_enabled=False, render_mode='rgb_array')[source]#
Every setting an env is built from; the defaults are the official values.
controller=Nonebuilds a floating-base env with no lower-body controller (fast, but not the benchmark substrate).episode_length=Nonetakes the task’s registered budget.- Parameters:
robot_model (Literal['g1_dex1'])
controller (ControllerConfig | None)
episode_length (int | None)
demo_down_sample_rate (int)
success_hold_seconds (float)
reach_tolerance (float | None)
initialization_profile (Literal['upstream', 'g1_id_v1'])
enable_all_floating_dof (bool)
control_pelvis (bool)
action_mode (Literal['absolute', 'delta'])
event_reward_shaping_enabled (bool)
event_reward_progress_scale (float)
event_reward_holding_bonus (float)
event_reward_lift_bonus (float)
event_reward_target_bonus (float)
wholebody (WholeBodyConfig | None)
frame_stack (int)
normalize_low_dim_obs (bool)
action_representation (Literal['absolute', 'upper_delta'])
upper_delta_scale_rad (float | None)
event_progress_enabled (bool)
render_mode (str)
- override(overrides=None, /, **kwargs)[source]#
Return a copy with the named fields replaced.
Keys are EnvConfig field names; an unknown key raises ValueError.
controller/wholebodyaccept None, a config instance, or a mapping of its fields merged into the current one.
- class bigym.loco.config.ControllerConfig(backend='groot_wbc_g1', base_action_mode='lowerbody_cmd', pitch_command=True, reset_warmup_steps=200, deterministic_reset=True, init_stance='keyframe', passive_base_tilt=True, default_height_cmd=0.74, init_pelvis_z=0.74, height_cmd_min=None, height_cmd_max=None, pitch_cmd_min=None, pitch_cmd_max=None, default_pitch_cmd=0.0, cmd_clip=1.0, wz_clip=1.0, use_height_cmd=True, skyhook_kp=0.0, skyhook_damping=0.0, model_path=None)[source]#
The lower-body controller in the loop (GR00T-WBC on the G1).
Nonecommand bounds mean “the ranges the backend declares”.- Parameters:
backend (str)
base_action_mode (Literal['lowerbody_cmd', 'legacy_delta'])
pitch_command (bool)
reset_warmup_steps (int | None)
deterministic_reset (bool)
init_stance (Literal['keyframe'])
passive_base_tilt (bool)
default_height_cmd (float)
init_pelvis_z (float)
height_cmd_min (float | None)
height_cmd_max (float | None)
pitch_cmd_min (float | None)
pitch_cmd_max (float | None)
default_pitch_cmd (float)
cmd_clip (float)
wz_clip (float)
use_height_cmd (bool)
skyhook_kp (float)
skyhook_damping (float)
model_path (str | None)
- class bigym.loco.config.WholeBodyConfig(mode='hierarchical_current', preserve_leg_proprio=False)[source]#
Outer action over the leg joints too (research modes, off by default).
- bigym.loco.config.resolve_config(task_name, config=None, overrides=None)[source]#
The config
makebuildstask_namefrom.configis None (the task’s official config) or anEnvConfigused as the base.overridesgo on top. A Noneepisode_lengthbecomes the task’s budget.
Task registry#
- class bigym.loco.tasks.TaskSpec(env_cls, episode_length, data_derived=False, overrides=<factory>)[source]#
One registered task: its class, budget and official-config differences.
env_clsis aBiGymEnvsubclass; the env builds it with theBiGymEnvconstructor arguments.- Parameters:
- bigym.loco.tasks.budget_provenance(name)[source]#
data_derivedwhen the budget came from demos, elseupstream_placeholder.
Contract layer#
Typed lower-body command declarations.
A lower-body backend consumes a command every control step. Every shipped
backend is velocity-conditioned (VELOCITY): the command is a flat vector
of scalar channels such as vx / vy / wz / height / torso_pitch. The
kind tag leaves room for other kinds of controller (SE3 end-effector
targets, reference-motion trackers) without changing the velocity backends.
CommandSpecis data: adapters declare their spec. Its bounds are the commands the backend accepts; GR00T-WBC publishes no training ranges, so its adapter declares benchmark clips.height/torso_pitch: the outer action-space bounds derive from the spec unless the config sets them (LowerBody.resolve_command_bounds; the config wins, so recorded demo action spaces do not move).vx/vy/wz: the env clips them with the symmetriccmd_clip/wz_clip; the spec bounds are not read. Per-channel spec bounds would change the action semantics (a substrate bump).rate: an optional slew limit in units/s,Nonefor a policy trained on step commands. It is declarative: adapters own their slew logic.
- class bigym.loco.command.CommandKind(*values)[source]#
What a lower-body command vector means.
- VELOCITY = 'velocity'#
Flat scalar channels, e.g. [vx, vy, wz, height, torso_pitch…].
groot_wbc_g1is this kind.
- EE_POSE = 'ee_pose'#
Reserved for SE3 hand/foot targets of whole-body controllers.
- MOTION_REF = 'motion_ref'#
Reserved for reference-motion trajectories (SONIC-style trackers).
- class bigym.loco.command.CommandField(name, unit, low, high, rate=None)[source]#
One scalar channel of a VELOCITY command.
- class bigym.loco.command.CommandSpec(kind, fields)[source]#
A backend’s full command declaration.
- Parameters:
kind (CommandKind)
fields (tuple[CommandField, ...])
- bigym.loco.command.velocity_spec(*, vx, vy, wz, height=None, height_rate=None, torso_pitch=None, torso_pitch_rate=None)[source]#
Build the common twist(+height)(+torso_pitch) VELOCITY spec.
The lower-body controller contract (thin required core).
LowerBodyController is the entire surface the BiGym env integration
relies on. Anything else a concrete adapter exposes is an implementation
detail. New backends normally subclass bigym.loco.base.LowerBodyBase
(the fat base class that absorbs joint addressing, reset pose application,
failure detection and declarative replay state) and provide a policy loader,
an obs builder and a policy step — a few hundred lines for a real backend —
but any object satisfying this protocol plugs in.
Contract summary:
controlled_jointsjoints whose position targets this backend owns.command_spectyped command declaration (see loco.command).control_dtseconds betweenstep()calls (1/control_hz).output_specnames + bounds of the produced targets.reset()re-anchor to the backend’s init pose, clear state.set_command(...)latch the current command (VELOCITY kind channels).step()run the policy once -> joint position targets, incontrolled_jointsorder.
is_failed()has the robot fallen / left the recoverable region.get_state()/set_state()every mutable field that shapes futuretargets, for bit-exact mid-episode save/restore (demo replay). MuJoCo qpos/qvel/ctrl/qacc_warmstart are snapshotted by the caller; this covers the controller-side remainder (obs histories, last actions, rate-limiter anchors, gait clocks…).
- class bigym.loco.controller.OutputSpec(joint_names, low, high)[source]#
Names and bounds of the joint position targets a backend produces.
- class bigym.loco.controller.LowerBodyController(*args, **kwargs)[source]#
Thin required core every lower-body backend implements.
- property controlled_joints: tuple[str, ...]#
Joints whose position targets this backend owns (incl. waist).
- property command_spec: CommandSpec#
Typed command declaration and its bounds.
- property output_spec: OutputSpec#
Names/bounds of the produced targets.
- set_command(cmd_vx, cmd_vy, cmd_wz, *, height=None, torso_pitch=None)[source]#
Latch the VELOCITY-kind command channels.
Keyword names match the
command_specfield names (height/torso_pitch). Channels absent fromcommand_specare ignored. Non-velocity command kinds (EE_POSE / MOTION_REF) will extend this surface as a pure addition; velocity backends stay untouched.
- class bigym.loco.base.LowerBodyBase[source]#
Shared pipeline for lower-body backends.
Subclasses set in
__init__:controlled_joints,command_spec,control_dt,controlled_range_low/controlled_range_high(frombuild_joint_ranges()),velocity_clip/yaw_rate_clip, and the latchedcommand,height_commandandlast_action.- command_spec: CommandSpec#
The backend’s typed command declaration and its bounds.
- property output_spec: OutputSpec#
Names and bounds of the joint targets produced by
step().
- set_command(cmd_vx, cmd_vy, cmd_wz, *, height=None, torso_pitch=None)[source]#
Default clip-on-set semantics (groot_wbc family).
Keyword names match the
command_specfield names.
- build_joint_ranges(joint_names)[source]#
The lower and upper limits of
joint_names; unlimited joints get +-inf.
- find_sensor(sensor_name, *, dim)[source]#
The
sensordataaddress of thedim-wide sensorsensor_name, or None.
- apply_pose(qpos_addresses, dof_addresses, targets, joint_names)[source]#
Write a joint pose + matching actuator targets and re-forward.
- get_state()[source]#
Snapshot every attribute named in
STATEFUL.Scalars become 0-d float32 arrays, arrays are float32 copies — matching the recorded snapshot format so demo npz files stay interchangeable.
Backends#
Lower-body backends: the registry and the built-in groot_wbc_g1.
BACKENDS maps each registered backend name to its
BackendBinding; register_backend() adds one. A backend name
with a colon ("pkg.module:ATTR") instead names a binding in an
importable module, imported on first use.
The vendored policy runtime needs only onnxruntime (a core dependency), imported when a controller is built, so the bigym core stays torch-free.
- bigym.loco.adapters.register_backend(name, binding)[source]#
Register
bindingundername(controller={"backend": name}).A name can be registered once; registering it again (
groot_wbc_g1included) raises ValueError.controller={"backend": "pkg.module:ATTR"}uses a binding without registering it.- Parameters:
name (str)
binding (BackendBinding)
- Return type:
None
- bigym.loco.adapters.resolve_backend_name(name)[source]#
Validate a backend name: registered, or an importable binding reference.
- class bigym.loco.BackendBinding[source]#
A lower-body backend’s hooks into env construction.
The env calls the hooks in this order:
robot_cls()while building the task env, thenbuild_controller()andconfigure_model()on the built env.envis the task’sBiGymEnv; a controller may rely onenv.model,env.data,env.robotandenv.action_space.configis the env’sControllerConfig: a backend reads the fields that apply to it and may ignore the rest.Subclass it and set
robot_models; overridebuild_controller(). The controller implementsLowerBodyController, most easily by subclassingLowerBodyBase.- supports_passive_base_tilt: bool = False#
True when the backend can run with the pelvis roll/pitch left passive (
ControllerConfig.passive_base_tilt).
- robot_cls(robot_cls, config)[source]#
The robot class to build the task env with.
robot_clsis the robot model’s floating-base class. Return it (the default), or a variant with the joints the controller drives actuated and their PD gains set (G1Dex1.variant).- Parameters:
robot_cls (type[Robot])
config (ControllerConfig)
- Return type:
type[Robot]
- build_controller(env, config, *, control_dt)[source]#
Build the controller on the freshly built task env.
control_dtis the env’s control step in seconds.- Parameters:
env (BiGymEnv)
config (ControllerConfig)
control_dt (float)
- Return type:
Demos#
Native demo schema v1 (frozen).
A native demo is one teleoperated episode recorded against the
controller-in-the-loop env (bigym.loco.env), stored as one .npz per
episode plus one metadata.json per collection directory. This module
freezes the contract those files satisfy so training/replay code can
validate instead of assuming.
Episode npz keys#
Per-step arrays, first axis = outer control steps (T):
rgb_obs(T, num_cameras, 3, H, W) uint8low_dim_obs(T, low_dim) float32 — raw (unstacked) layout, seelow_dim_component_slices()of the env for the component mapaction(T, action_dim) float32 — the NORMALIZED outer action in [-1, 1] (commands + upper-body targets + grippers), exactly what the collector fed env.step(). Decode to physical units via the per-dimaction_statsmin/max recorded in metadata.json (the collection envelope;env.get_demos()adopts it at load). The RAW per-step values live in the diagnosticraw_outer_actionarray, not here.reward/discount/demo/is_expert(T, 1) float32event_progress(T, 1) float32 — optional
Per-episode scalars / snapshots:
seed(1,) int64 — env reset seedpre_engage_steps(1,) int64 — steps between reset and VR engageengage-state snapshot:
init_qpos,init_qvel,init_ctrl,init_qacc_warmstart(andinit_actwhen present) — MuJoCo state at the engage momentlb_state.*— the flattened lower-body state snapshot (env.get_lowerbody_state():ctrl.*controller fields per LowerBodyController.get_state(),env.*env-side fields). Restoring MuJoCo state +lb_state.*reproduces the episode bit-exactly on the Newton-pinned scenes.
metadata.json#
Written once per collection dir by the collector. Required keys:
format:"bigym_replay_npz"pipeline_version: collector pipeline version (string, e.g."2026-08-26-success-hold-training-view-v1")task: the flat env settings (must includeepisode_lengthanddemo_down_sample_rate— training MUST match these), plus the collection and training success holdslowerbody_policy: the lower-body controller settings, flatoptional
env_config: the fullEnvConfigthe batch was recorded in;EnvConfig.from_metadatareads it, and needs itreset_semantics: init keyframe / pelvis z / warmup stepsaction_semantics: robot model, backend, base_action_mode, stick mapoptional
action_stats: per-dim min/max for [-1,1] rescaling
- bigym.loco.demos.schema.validate_episode(episode)[source]#
Return a list of schema violations (empty = valid).
- bigym.loco.demos.schema.validate_metadata(metadata)[source]#
Return a list of metadata violations (empty = valid).
npz + metadata.json IO for native demos (schema: bigym.loco.demos.schema).
- bigym.loco.demos.io.load_episode(path)[source]#
Load one episode npz into a plain dict (all arrays materialized).
- bigym.loco.demos.io.save_episode(episode, path, *, validate=True)[source]#
Atomically save one episode npz (write-then-rename).
- bigym.loco.demos.io.load_metadata(demo_dir)[source]#
Load the collection dir’s metadata.json (None when absent).
Cut a raw replay-format demo batch to its training view.
An episode succeeds once the task predicate has held for
success_hold_seconds; the recorded episode ends on that frame with the
single terminal reward (1) and discount (0). The collector records with a
longer hold than training and evaluation use, so the raw batch is cut before
it trains: each episode is replayed from its engage snapshot on the task’s
official env and ends on the control step where that env latches success. A
predicate that held long enough, broke and restarted before the collection
hold completed makes that step fall anywhere before the raw end, so no fixed
trim finds it. The cut goes to a NEW directory and the source batch is never
modified. Usage:
python -m bigym.loco.demos.success_hold --demo-dir <batch>
- bigym.loco.demos.success_hold.latch_batch(demo_dir, out_dir=None)[source]#
Replay and cut a raw batch; see
latch_steps()andcut_batch().
Hugging Face Hub access to the BiGym 2.0 demonstration dataset.
The public demonstrations live in one Hugging Face dataset repository with one top-level folder per task, each a LeRobot v3 lossless export:
<repo>/
move_plate/ data/chunk-000/file-*.parquet + meta/ + metadata.json
reach_target_single/ ...
Two ways to get them onto a machine:
Lazy, per task:
task_dir()fetches one task’s folder the first time it is needed.env.get_demos()calls it, so running a task pulls exactly that task’s demonstrations (a few hundred MB to a few GB).Ahead of time:
bigym-download --all(ordownload_all()) mirrors the whole dataset;bigym-download --task move_plate pick_boxfetches a subset.--local-dir PATHwrites the files into a plain folder instead of the cache, whichbigym-view --demo-dir PATHopens.
Both go through huggingface_hub.snapshot_download and share its cache
($HF_HOME / $HF_HUB_CACHE; default ~/.cache/huggingface), so a
pre-download and a later lazy load never fetch a file twice, and the usual
Hub knobs apply (HF_TOKEN for private/gated repos, HF_HUB_OFFLINE=1
to refuse network access and use the cache only).
The repository defaults to DEFAULT_DATASET_REPO; override it with the
BIGYM_DATASET_REPO environment variable (and optionally pin a revision
with BIGYM_DATASET_REVISION).
The dataset repository has no demonstrations for the requested task.
For a benchmark task this means its demonstrations have not been published yet: the dataset is released task by task, and the message lists what is published and what is still pending.
Read a LeRobot v3 lossless task export back into replay-format episodes.
This is the reader half of bigym.loco.demos.lerobot_export. It reads
the on-disk format directly (pyarrow + PNG decode); the lerobot package
is never imported, so it works on every Python the core package supports.
Only lossless PNG-mode datasets are accepted: video-mode cameras are lossy and would silently change training inputs.
Alignment: the export (meta/alignment.json, version 2) stores
transition-scoped features shifted so LeRobot frame k pairs obs[k] with the
action executed FROM it; this reader undoes the shift (replay index 0 comes
back from the first_transition sidecar, the repeated final frame is
dropped), so the episodes come back exactly as the collector wrote them:
row t holds the observation at step t together with the action, reward and
discount of the transition that PRODUCED it (row 0 is the reset row with a
zero action).
- bigym.loco.demos.dataset.load_episodes(task_dir, max_episodes=-1, *, action_representation='absolute', upper_delta_scale_rad=None)[source]#
Reconstruct replay-format episode dicts from one task’s LeRobot export.
Returns
(source_name, episode)pairs in dataset order. Each episode maps feature names to[T, ...]arrays:rgb_obs[T, cams, 3, H, W]uint8 (camera order from the collector metadata),low_dim_obs[T, D],action[T, A](normalized outer action),reward,discount,demo,is_expert,event_progress[T, 1]plus any per-step extras the collector stored.With the default
action_representation="absolute"the arrays are the collector’s bit for bit."upper_delta"derives a training view in memory; the dataset is never modified.
Demo collection#
- class bigym.vr.collect.config.CollectConfig(task, out_dir=None, episodes=60, seed=1, collect_success_hold_seconds=3.0, episode_steps=None, episode_seconds=None, max_steps=None, keep_empty_session=False, keep_black_rgb=False, store_event_progress=True, store_fullbody=True, yaw_mode='base', base_vx_scale=0.35, base_vy_scale=0.25, base_wz_scale=0.5, base_z_scale=0.004, height_cmd_max=0.8, pitch_rate=0.8, stick_deadzone=0.15, base_cmd_slew=0.7, recenter_height_offset=0.0, vr_space_mode='follow_head', settle_view='curtain', hud='minimal', resolution='lq', mujoco_gl='glfw', openxr_log_level='error', spectator='none', spectator_hz=15.0, spectator_port=8080, spectator_gpu=None, frame_spike_ms=20.0, operator=None, export_lerobot=False, task_text=None)[source]#
One VR collection session (bigym-collect); each field is a flag.
- Parameters:
task (str)
out_dir (Path | None)
episodes (int)
seed (int)
collect_success_hold_seconds (float)
episode_steps (int | None)
episode_seconds (float | None)
max_steps (int | None)
keep_empty_session (bool)
keep_black_rgb (bool)
store_event_progress (bool)
store_fullbody (bool)
yaw_mode (Literal['base', 'none'])
base_vx_scale (float)
base_vy_scale (float)
base_wz_scale (float)
base_z_scale (float)
height_cmd_max (float)
pitch_rate (float)
stick_deadzone (float)
base_cmd_slew (float | None)
recenter_height_offset (float)
vr_space_mode (Literal['follow_head', 'fixed'])
settle_view (Literal['curtain', 'live'])
hud (Literal['minimal', 'full', 'off'])
resolution (Literal['lq', 'mq', 'hq'])
mujoco_gl (Literal['glfw', 'egl', 'osmesa'])
openxr_log_level (Literal['trace', 'debug', 'info', 'warn', 'error'])
spectator (Literal['none', 'mujoco', 'viser', 'both'])
spectator_hz (float)
spectator_port (int)
spectator_gpu (int | None)
frame_spike_ms (float)
operator (str | None)
export_lerobot (bool)
task_text (str | None)
Evaluation#
- bigym.loco.eval.runner.evaluate(policy, *, task_name, method, checkpoint='n/a', episodes=100, config=None, overrides=None, env=None, verbose=True)[source]#
Run one evaluation block; return summary (+ leaderboard record).
Pass either
config/overrides(forwarded tobigym.loco.make()withtask_name; the env is created and closed here) or a readyenv(kept open).Returns
{"task", "method", "episodes", "success_rate", "episode_rewards", "substrate", "env_config", "protocol_violations", ["record"]};record(the protocol-v1 leaderboard entry) only for a full block on an official env.
The frozen BiGym 2.0 evaluation protocol.
One definition shared by every leaderboard entry:
100 episodes per (task, checkpoint) evaluation;
fixed seed block starting at 620000, reserved for evaluation by convention (episode i uses
seed_base + i; training-seed overlap is not checked here);checkpoints every 5k training steps, each evaluated on the same block; the reported number is the MEAN OF THE LAST 5 CHECKPOINTS (80k..100k of a 101k budget) +- standard error.
peak_final_selectionserves the appendix tables. Selecting the maximum on the same test block introduces selection bias; do not use it as the main-table result or treat correlated checkpoints as independent training seeds. This module records individual checkpoints, not that aggregate; disclose checkpoint/seed/episode aggregation when reporting it;an episode is a success when the task reported success and the robot never fell (
is_success()). Evaluation must keep reward shaping off.
- bigym.loco.eval.protocol.is_success(env)[source]#
Whether the env’s episode so far counts as a success.
- Parameters:
env (Any) – An env from
bigym.loco.make()ormake_gym, or any wrapper that forwardsepisode_succeededandepisode_fell.- Returns:
True when the task reported success (
env.episode_succeeded()) and the robot never fell (env.episode_fell()).- Return type:
- bigym.loco.eval.protocol.eval_seeds(episodes=100, *, seed_base=620000)[source]#
The per-episode reset seeds of one evaluation block.
- bigym.loco.eval.protocol.peak_final_selection(step_to_score)[source]#
Pick the peak-score and final checkpoints from step->score curves.
- bigym.loco.eval.protocol.protocol_violations(env)[source]#
Why an env’s results are not official; empty when they are.
Compares the configuration the env was built from (
env.config) with its task’s official configuration (env.official_config). Every result-affecting field that differs is listed; fields that only change rendering or the policy-side view of observations and actions are exempt (seebigym.loco.config.affects_results()).- Parameters:
env (Any) – An env from
bigym.loco.make()ormake_gym, or any wrapper that forwardsconfigandofficial_config.- Returns:
One
"field: official X, env Y"line per departing field.- Return type:
Coding-agent benchmark#
The coding-agent benchmark: an LLM agent writes the control policy.
An agent works in a sandbox directory containing a task brief, a policy
template and a harness that talks to an environment server over a unix socket.
The server (bigym.loco.agent.server) owns the simulator, counts the
agent’s interaction budget, keeps a ledger and refuses the hidden evaluation
seeds. When the session ends, bigym.loco.agent.evaluate scores the
submitted policy.py on those hidden seeds with the same episode loop the
agent developed against (bigym.loco.agent.episode), the policy running
in an interpreter of its own that holds no simulator
(bigym.loco.agent.policy_process).
Public surface:
from bigym.loco.agent import EnvTools, EnvToolsConfig, make_env, Tools
from bigym.loco.agent import load_policy, run_episode
EnvTools(task, EnvToolsConfig(...)) builds the environment a policy runs
against, with the settings the command line takes.
Names are resolved lazily, so importing this package pulls in neither the
simulator nor the optional bigym[agent] dependencies.
- bigym.loco.agent.envtools.make_env(task, pitch=None)[source]#
Build the benchmark environment for a task.
- Parameters:
- Returns:
The controller-in-the-loop env from
bigym.loco.make().
- class bigym.loco.agent.episode.Tools(env)[source]#
What a policy may call besides reading observations (no reset/step).
- bigym.loco.agent.episode.load_policy(path)[source]#
Import
pathand instantiate thePolicyclass it defines.- Parameters:
path (Path) – Path to the policy file.
- Returns:
A fresh
Policyinstance.- Raises:
AttributeError – The file defines no
Policyclass.