Getting started#
Install BiGym 2.0 first.
Make and reset#
make(task) builds the task’s official configuration: the Unitree G1 with
Dex1-1 two-finger grippers (g1_dex1), the GR00T-WBC lower body
(groot_wbc_g1), three 84×84 cameras, and the task’s episode budget and
success hold.
from bigym.loco import make
env = make("move_plate")
timestep = env.reset(seed=620000)
env.config # the resolved EnvConfig
reset runs a warmup before it returns: the lower-body controller steps
with the arms held while the robot settles into its stance. Your policy
takes over after the warmup. reset_warmup_steps defaults to 200 control
steps (4 s). Never set it to 0. The demonstrations start after the warmup,
so evaluation has to warm up too.
One step#
make returns a dm_env-style wrapper. step(action) advances one 50 Hz
control step and returns an ExtendedTimeStep with rgb_obs,
low_dim_obs, reward and step_type.
import numpy as np
spec = env.action_spec() # shape (20,), float32
action = np.zeros(spec.shape, dtype=spec.dtype)
timestep = env.step(action)
timestep.rgb_obs.shape # (3, 3, 84, 84): head, right_wrist, left_wrist
timestep.low_dim_obs.shape # (50,)
timestep.step_type.name # 'MID'
Actions are normalized to [-1, 1] per slot, and the env maps each slot onto its range. A zero action commands zero base velocity and the middle of every other range.
What am I controlling?#
env.wholebody_action_layout()
# {'action_dim': 20, 'base_dim': 4, 'limb_names': (...14 arm joints...),
# 'gripper_count': 2, 'action_low': ..., 'action_high': ..., ...}
env.low_dim_component_slices()
# {'proprioception': (0, 44), 'proprioception_grippers': (44, 46),
# 'proprioception_floating_base': (46, 50)}
The four base slots are [vx, vy, height, wz]: forward and lateral velocity
in m/s, absolute pelvis height in m and yaw rate in rad/s. The 14 arm slots
are absolute joint position targets. The 2 gripper slots open (0) or close
(1) each hand.
Those are the move_plate values. The 14 tasks with the torso-pitch command
add a fifth base slot and have a 21-dim action and a 56-dim low-dimensional
observation. Read the layout from the env.
Official configuration has the
full tables.
Overrides#
Keyword overrides or an EnvConfig change any setting. The env records what
departs from the official configuration, and evaluation reports a run as
unofficial when a departure affects results:
env = make("move_plate", camera_keys=("head",), controller={"cmd_clip": 0.5})
env.config_overrides # {'controller.cmd_clip': 0.5, 'camera_keys': ('head',)}
env = make("move_plate", env.config.override(camera_shape=(128, 128)))
Demonstrations#
from bigym.loco import make_gym
env = make_gym("move_plate")
demos = env.get_demos(60) # downloads move_plate's demos on first use
demo = demos[0]
demo["obs"]["rgb"].shape # (T + 1, 3, 3, 84, 84)
demo["action"].shape # (T, 20), in env.action_space
make_gym returns a gymnasium env with the same task.
bigym-download --all fetches every task ahead of time. See
Demonstrations.
Check the install#
MUJOCO_GL=egl uv run python examples/loco_random_agent.py
It ends with one line per seed, such as
seed 620000: total reward 0.000, last step MID.