File size: 19,339 Bytes
700dd75 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 | # Training Code Structure
This page describes the Python training codebase under `gear_sonic/`, covering directory layout, the training pipeline, configuration system, key modules, and evaluation scripts.
---
## Directory Layout
```
gear_sonic/
βββ train_agent_trl.py # Main training entry point
βββ eval_agent_trl.py # Single-checkpoint evaluation
βββ eval_exp.py # Checkpoint monitor (continuous eval)
βββ config/ # Hydra configuration hierarchy
β βββ base.yaml # Global defaults (seed, num_envs, paths)
β βββ base_eval.yaml # Eval-specific global defaults
β βββ eval_exp.yaml # Checkpoint monitor config
β βββ base/ # Hydra plumbing (output dirs, resolvers)
β βββ algo/ # PPO hyperparameters
β βββ actor_critic/ # Actor-critic architecture configs
β β βββ encoders/ # Per-encoder MLP configs (g1, smpl, teleop)
β β βββ decoders/ # Decoder MLP configs (g1_kin, g1_dyn)
β β βββ critics/ # Critic backbone configs
β β βββ quantizers/ # FSQ quantizer config
β β βββ universal_token/ # Assembled encoder+decoder+quantizer presets
β βββ aux_losses/ # Auxiliary loss definitions
β βββ callbacks/ # Training callback configs
β βββ exp/ # Experiment presets (compose all pieces)
β βββ manager_env/ # Environment MDP component configs
β βββ opt/ # Logging options (wandb)
β βββ trainer/ # Trainer class selection
βββ envs/ # IsaacLab environment wrappers
β βββ manager_env/
β β βββ modular_tracking_env_cfg.py # Scene, sensors, robot articulation
β β βββ robots/ # Per-robot configs (g1.py, h2.py)
β β βββ mdp/ # MDP components (see below)
β βββ wrapper/
β β βββ manager_env_wrapper.py # RL-facing env wrapper
β βββ env_utils/ # Joint ordering utilities
βββ trl/ # Training modules (PPO, actor-critic, losses)
β βββ trainer/
β β βββ ppo_trainer.py # Base PPO trainer
β β βββ ppo_trainer_aux_loss.py # PPO + auxiliary losses (SONIC)
β βββ modules/
β β βββ actor_critic_modules.py # Actor, Critic classes
β β βββ universal_token_modules.py # UniversalTokenModule (SONIC ATM)
β β βββ base_module.py # Shared MLP building blocks
β β βββ data_utils.py # Batch/data helpers
β βββ losses/
β β βββ token_losses.py # Reconstruction & latent auxiliary losses
β βββ callbacks/ # Runtime callbacks
β β βββ im_eval_callback.py # Imitation evaluation metrics
β β βββ im_resample_callback.py # Adaptive motion resampling
β β βββ model_save_callback.py # Checkpoint saving
β β βββ wandb_callback.py # W&B logging
β β βββ read_eval_callback.py # Read eval results from disk
β βββ utils/ # Math, rotation, scheduling utilities
βββ utils/ # Shared utilities
β βββ motion_lib/ # Motion library loading (PKL format)
β βββ mujoco_sim/ # MuJoCo sim-to-sim bridge
β βββ teleop/ # VR teleoperation helpers
βββ data/ # Robot models, URDF/USD assets
βββ data_process/ # Motion data conversion scripts
βββ scripts/ # MuJoCo sim loop, misc tools
```
---
## Training Pipeline
Running `python gear_sonic/train_agent_trl.py +exp=manager/universal_token/all_modes/sonic_release` executes the following steps:
### 1. Configuration Loading
The entry point uses `@hydra.main(config_path="config", config_name="base")`. The `+exp=...` argument selects an experiment preset that composes all sub-configs:
```
base.yaml # Global defaults
βββ +exp=manager/universal_token/all_modes/sonic_release
βββ /algo: ppo_im_phc # PPO hyperparameters
βββ /actor_critic: universal_token/all_mlp_v1
β βββ encoders/g1_mf_mlp, smpl_mlp, teleop_mlp
β βββ decoders/g1_kin_mf_mlp, g1_dyn_mlp
β βββ quantizers/fsq
β βββ critics/mlp
βββ /manager_env: base_env # Environment config
β βββ observations/{tokenizer, policy, critic}
β βββ rewards/tracking/base_5point_local_feet_acc
β βββ terminations/tracking/base_adaptive_strict_ori_foot_xyz
β βββ events/tracking/level0_4
βββ /aux_losses: universal_token/g1_recon_and_all_latent
βββ /trainer: trl_ppo_aux
βββ /callbacks: model_save, wandb, read_eval, im_resample
```
### 2. Simulator and Accelerator Init
After config resolution, the script:
1. Parses TRL `PPOConfig` / `ScriptArguments` / `ModelConfig` from the config dict.
2. Creates a HuggingFace `Accelerator` for multi-GPU support (DDP).
3. Launches the IsaacLab `AppLauncher` to start the Isaac Sim runtime.
4. Saves `config.yaml` and `meta.yaml` to the experiment directory.
### 3. Environment Creation
`create_manager_env()` instantiates the IsaacLab `ManagerBasedRLEnv` from the composed environment config, then wraps it with `ManagerEnvWrapper`:
```
ManagerBasedRLEnv (IsaacLab)
βββ ManagerEnvWrapper
βββ Observation spaces (policy, critic, tokenizer groups)
βββ Motion command manager (motion_lib)
βββ Action transform module (optional, for pretrained ATM)
βββ Keyboard / visualization hooks
```
### 4. Policy and Value Model Creation
The actor and critic are instantiated from the algo config. For SONIC training, the actor backbone is `UniversalTokenModule`:
```python
# Simplified from train_agent_trl.py
policy = custom_instantiate(config.algo.config.actor, env_config=env.config, ...)
value_model = custom_instantiate(config.algo.config.critic, env_config=env.config, ...)
```
The `Actor` wraps `UniversalTokenModule` as its backbone and adds a diagonal Gaussian distribution for exploration. The `Critic` wraps a separate MLP backbone.
### 5. PPO Training Loop
The `TRLAuxLossPPOTrainer.train()` method runs the main loop:
```
for iteration in range(num_learning_iterations):
# 1. Rollout: collect num_steps_per_env transitions
for step in range(num_steps_per_env):
actions = policy.rollout(obs_dict)
obs_dict, rewards, dones, infos = env.step(actions)
store(obs, actions, rewards, values, log_probs)
# 2. GAE: compute advantages and returns
advantages = generalized_advantage_estimation(rewards, values, dones)
# 3. PPO update: num_ppo_epochs over mini-batches
for epoch in range(num_ppo_epochs):
for mini_batch in shuffle_and_split(rollout_data):
policy_loss = clipped_surrogate_objective(...)
value_loss = clipped_value_loss(...)
aux_loss = sum(coef_i * aux_loss_i) # encoder reconstruction, etc.
total_loss = policy_loss + value_loss_coef * value_loss
+ aux_loss_scale * aux_loss
optimizer.step(total_loss)
# 4. Post-update: sync running stats, adaptive sampling, callbacks
update_scheduled_params(...) # learning rate, domain randomization
callbacks.on_step_end(...) # checkpointing, evaluation, logging
```
---
## Configuration System
The configuration system uses [Hydra](https://hydra.cc/) with config groups and composition.
### Hierarchy
| Level | Path | Purpose |
|---|---|---|
| **Global** | `config/base.yaml` | Seed, num_envs, paths, wandb toggle |
| **Algorithm** | `config/algo/ppo_im_phc.yaml` | PPO hyperparameters, learning rates, epochs |
| **Actor-Critic** | `config/actor_critic/` | Network architecture (encoders, decoders, critic) |
| **Environment** | `config/manager_env/` | Observations, rewards, terminations, events |
| **Auxiliary Losses** | `config/aux_losses/` | Reconstruction and latent alignment losses |
| **Trainer** | `config/trainer/` | Trainer class selection (PPO or PPO+AuxLoss) |
| **Callbacks** | `config/callbacks/` | Checkpointing, evaluation, W&B logging |
| **Experiment** | `config/exp/` | Preset that composes all the above |
### Experiment Presets
Experiment configs live under `config/exp/` and use the `@package _global_` directive to set values at the root level. They compose all component configs via `defaults`:
```yaml
# config/exp/manager/universal_token/all_modes/sonic_release.yaml
defaults:
- /algo: ppo_im_phc
- /manager_env: base_env
- override /actor_critic: universal_token/all_mlp_v1
- override /manager_env/observations/tokenizer: unitoken_all_noz
- override /manager_env/observations/policy: local_dir_hist
- override /manager_env/rewards: tracking/base_5point_local_feet_acc
- override /manager_env/terminations: tracking/base_adaptive_strict_ori_foot_xyz
- override /manager_env/events: tracking/level0_4
# ...
```
### Key Config Parameters
| Parameter | Default | Description |
|---|---|---|
| `num_envs` | 4096 | Number of parallel simulation environments |
| `algo.config.num_learning_iterations` | 100000 | Total training iterations |
| `algo.config.num_steps_per_env` | 32 | Rollout horizon per iteration |
| `algo.config.num_learning_epochs` | 5 | PPO epochs per iteration |
| `algo.config.num_mini_batches` | 4 | Mini-batches per PPO epoch |
| `algo.config.actor_learning_rate` | 2e-5 | Actor learning rate |
| `algo.config.critic_learning_rate` | 1e-3 | Critic learning rate |
| `algo.config.clip_param` | 0.2 | PPO clipping parameter |
| `algo.config.init_noise_std` | 0.05 | Initial exploration noise std |
| `algo.config.save_interval` | 500 | Checkpoint save frequency (iterations) |
---
## Universal Token Module
The `UniversalTokenModule` implements SONIC's action transform module (ATM) -- the core architecture that maps diverse motion inputs into a shared token space.
### Architecture
```
βββββββββββββββ
G1 obs ββββΊ β G1 Encoder ββββ
βββββββββββββββ β
βββββββββββββββ β βββββββββββ βββββββββββββββ
Teleop obsββββΊ βTeleop Encdr ββββΌβββΊ β FSQ ββββΊ β G1 Dynamic ββββΊ joint actions
βββββββββββββββ β βQuantizerβ β Decoder β
βββββββββββββββ β βββββββββββ βββββββββββββββ
SMPL obs ββββΊ β SMPL Encoderββββ β
βββββββββββββββ β βββββββββββββββ
βββββββββΊ βG1 Kinematic ββββΊ (aux loss only)
β Decoder β
βββββββββββββββ
```
**Encoders** map different observation modalities into a shared latent space. Each encoder is an MLP that takes modality-specific tokenizer observations and outputs a fixed-size latent vector. During training, one encoder is sampled per environment according to `encoder_sample_probs`.
**FSQ Quantizer** discretizes the continuous latent into a finite set of tokens using Finite Scalar Quantization. Each latent dimension is independently quantized to one of `fsq_level_list` discrete levels. This produces a compact, discrete token representation.
**Decoders** reconstruct outputs from the quantized tokens plus proprioception:
- **G1 Dynamic Decoder** (`g1_dyn`): Produces joint-space actions fed to the actuators. This is the only decoder used at deployment time.
- **G1 Kinematic Decoder** (`g1_kin`): Reconstructs future motion frames from tokens. Used only during training to compute reconstruction auxiliary losses.
### Latent Residual Mode
For downstream tasks (e.g., object manipulation), an external policy can inject corrections into the token space without retraining the base ATM:
| Mode | Behavior |
|---|---|
| `post_quantization` (default) | Residual added after FSQ quantization |
| `pre_quantization` | Residual added before FSQ; the sum gets quantized |
| `pre_quantization_replace` | Latent is replaced entirely by the residual |
### Encoder Sampling
During training, each environment is randomly assigned an encoder per episode according to `encoder_sample_probs`. The `encoder_index` observation tells the module which encoder produced the current token. At deployment, only one encoder is active (selected by the observation configuration).
---
## Environment Structure
The training environment is built on IsaacLab's `ManagerBasedRLEnv` and uses a modular MDP design where each component is configured independently via YAML.
### MDP Components
All MDP components live in `gear_sonic/envs/manager_env/mdp/`:
| Module | Config path | Description |
|---|---|---|
| `observations.py` | `config/manager_env/observations/` | Observation terms for policy, critic, and tokenizer groups |
| `actions.py` | `config/manager_env/actions/` | Joint position action space |
| `rewards.py` | `config/manager_env/rewards/` | Reward terms (tracking, regularization) |
| `terminations.py` | `config/manager_env/terminations/` | Episode termination conditions |
| `events.py` | `config/manager_env/events/` | Domain randomization events |
| `commands.py` | `config/manager_env/commands/` | Motion command generation (motion library) |
| `curriculum.py` | `config/manager_env/curriculum/` | Curriculum schedules |
| `terrain.py` | (inline) | Terrain generation |
| `recorders.py` | `config/manager_env/recorders/` | Video recording |
### Observation Groups
Observations are split into groups, each with its own config file:
| Group | Purpose | Example terms |
|---|---|---|
| **policy** | Direct input to the policy MLP | joint_pos, joint_vel, base_ang_vel, gravity_dir, last_actions |
| **critic** | Privileged observations for the value function | All policy obs + base_lin_vel, body_pos, body_ori |
| **tokenizer** | Input to the UniversalTokenModule encoders | Multi-future joint commands, SMPL joints, VR targets, anchor orientations |
### Reward Terms
Reward configs compose individual terms from `config/manager_env/rewards/terms/`. Key tracking rewards:
| Term | Description |
|---|---|
| `tracking_relative_body_pos` | Track reference body positions (5-point: root, wrists, feet) |
| `tracking_relative_body_ori` | Track reference body orientations |
| `tracking_anchor_pos` | Track root anchor position |
| `tracking_anchor_ori` | Track root anchor orientation |
| `tracking_body_linvel` | Track reference body linear velocities |
| `tracking_body_angvel` | Track reference body angular velocities |
| `action_rate_l2` | Penalize action jerk |
| `feet_acc` | Penalize foot acceleration (smoothness) |
### ManagerEnvWrapper
`ManagerEnvWrapper` bridges the IsaacLab environment with the RL training loop. It handles:
- Flattening observation dicts for the policy
- Applying the optional pretrained action transform module
- Motion replay mode
- Debug visualization and keyboard controls
---
## Evaluation Scripts
### eval_agent_trl.py -- Single Checkpoint
Loads a single checkpoint and runs evaluation in Isaac Sim. Automatically reads the training `config.yaml` from the checkpoint directory to reconstruct the full configuration.
```bash
# Interactive visualization
python gear_sonic/eval_agent_trl.py +checkpoint=path/to/model.pt +headless=False ++num_envs=1
# Headless with video rendering
python gear_sonic/eval_agent_trl.py +checkpoint=path/to/model.pt +headless=True \
++num_envs=16 +run_once=True \
++manager_env.config.save_rendering_dir=path/to/output \
++manager_env.config.render_results=True \
+manager_env/recorders=render
```
Key features:
- Merges training config with eval overrides (`eval_overrides` in config)
- Removes train-only events and terminations automatically
- Supports `+run_once=True` to exit after all environments complete one episode
- Handles `+metrics_file` to render worst-performing motions from a prior eval
### eval_exp.py -- Checkpoint Monitor
`CheckpointEvaluator` continuously monitors an experiment directory for new checkpoints and evaluates them sequentially. It runs as a companion process alongside training.
```bash
python gear_sonic/eval_exp.py ++experiment_dir=path/to/experiment
```
For each new checkpoint, it:
1. Runs metrics evaluation (launches `eval_agent_trl.py` via subprocess)
2. Runs video rendering for the hardest motions
3. Logs results and videos to W&B (resuming the training run)
4. Marks each checkpoint as evaluated to avoid redundant work
Configuration (`config/eval_exp.yaml`):
| Parameter | Description |
|---|---|
| `experiment_dir` | Path to the training experiment directory |
| `scan_interval` | Seconds between checkpoint scans (default: 60) |
| `num_eval_envs` | Number of environments for metric evaluation |
| `num_render_videos` | Number of videos to render per checkpoint |
| `eval_frequency` | Only evaluate every N-th checkpoint (default: all) |
| `single_pass` | Evaluate pending checkpoints once and exit |
---
## Key Classes Reference
| Class | Module | Description |
|---|---|---|
| `Actor` | `trl/modules/actor_critic_modules.py` | Policy network: backbone + diagonal Gaussian. Maintains observation buffer for temporal models. |
| `Critic` | `trl/modules/actor_critic_modules.py` | Value function network: backbone + scalar output. Supports running mean/std normalization. |
| `UniversalTokenModule` | `trl/modules/universal_token_modules.py` | SONIC ATM: multi-encoder, FSQ quantizer, multi-decoder. Computes auxiliary reconstruction losses. |
| `TRLPPOTrainer` | `trl/trainer/ppo_trainer.py` | Base PPO trainer adapted from HuggingFace TRL. Handles rollout collection, GAE, and gradient updates. |
| `TRLAuxLossPPOTrainer` | `trl/trainer/ppo_trainer_aux_loss.py` | Extends `TRLPPOTrainer` with auxiliary loss support (reconstruction, latent alignment). |
| `PolicyAndValueWrapper` | `trl/trainer/ppo_trainer.py` | Wraps policy + value model into a single `nn.Module` for DDP-safe forward passes. |
| `ManagerEnvWrapper` | `envs/wrapper/manager_env_wrapper.py` | Bridges IsaacLab `ManagerBasedRLEnv` with the training loop. Handles obs flattening, action transforms, replay. |
| `CheckpointEvaluator` | `eval_exp.py` | Monitors experiment directory, evaluates new checkpoints, logs to W&B. |
|