File size: 19,339 Bytes
700dd75
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
# Training Code Structure

This page describes the Python training codebase under `gear_sonic/`, covering directory layout, the training pipeline, configuration system, key modules, and evaluation scripts.

---

## Directory Layout

```
gear_sonic/
β”œβ”€β”€ train_agent_trl.py          # Main training entry point
β”œβ”€β”€ eval_agent_trl.py           # Single-checkpoint evaluation
β”œβ”€β”€ eval_exp.py                 # Checkpoint monitor (continuous eval)
β”œβ”€β”€ config/                     # Hydra configuration hierarchy
β”‚   β”œβ”€β”€ base.yaml               # Global defaults (seed, num_envs, paths)
β”‚   β”œβ”€β”€ base_eval.yaml          # Eval-specific global defaults
β”‚   β”œβ”€β”€ eval_exp.yaml           # Checkpoint monitor config
β”‚   β”œβ”€β”€ base/                   # Hydra plumbing (output dirs, resolvers)
β”‚   β”œβ”€β”€ algo/                   # PPO hyperparameters
β”‚   β”œβ”€β”€ actor_critic/           # Actor-critic architecture configs
β”‚   β”‚   β”œβ”€β”€ encoders/           # Per-encoder MLP configs (g1, smpl, teleop)
β”‚   β”‚   β”œβ”€β”€ decoders/           # Decoder MLP configs (g1_kin, g1_dyn)
β”‚   β”‚   β”œβ”€β”€ critics/            # Critic backbone configs
β”‚   β”‚   β”œβ”€β”€ quantizers/         # FSQ quantizer config
β”‚   β”‚   └── universal_token/    # Assembled encoder+decoder+quantizer presets
β”‚   β”œβ”€β”€ aux_losses/             # Auxiliary loss definitions
β”‚   β”œβ”€β”€ callbacks/              # Training callback configs
β”‚   β”œβ”€β”€ exp/                    # Experiment presets (compose all pieces)
β”‚   β”œβ”€β”€ manager_env/            # Environment MDP component configs
β”‚   β”œβ”€β”€ opt/                    # Logging options (wandb)
β”‚   └── trainer/                # Trainer class selection
β”œβ”€β”€ envs/                       # IsaacLab environment wrappers
β”‚   β”œβ”€β”€ manager_env/
β”‚   β”‚   β”œβ”€β”€ modular_tracking_env_cfg.py   # Scene, sensors, robot articulation
β”‚   β”‚   β”œβ”€β”€ robots/             # Per-robot configs (g1.py, h2.py)
β”‚   β”‚   └── mdp/                # MDP components (see below)
β”‚   β”œβ”€β”€ wrapper/
β”‚   β”‚   └── manager_env_wrapper.py  # RL-facing env wrapper
β”‚   └── env_utils/              # Joint ordering utilities
β”œβ”€β”€ trl/                        # Training modules (PPO, actor-critic, losses)
β”‚   β”œβ”€β”€ trainer/
β”‚   β”‚   β”œβ”€β”€ ppo_trainer.py          # Base PPO trainer
β”‚   β”‚   └── ppo_trainer_aux_loss.py # PPO + auxiliary losses (SONIC)
β”‚   β”œβ”€β”€ modules/
β”‚   β”‚   β”œβ”€β”€ actor_critic_modules.py     # Actor, Critic classes
β”‚   β”‚   β”œβ”€β”€ universal_token_modules.py  # UniversalTokenModule (SONIC ATM)
β”‚   β”‚   β”œβ”€β”€ base_module.py              # Shared MLP building blocks
β”‚   β”‚   └── data_utils.py              # Batch/data helpers
β”‚   β”œβ”€β”€ losses/
β”‚   β”‚   └── token_losses.py     # Reconstruction & latent auxiliary losses
β”‚   β”œβ”€β”€ callbacks/              # Runtime callbacks
β”‚   β”‚   β”œβ”€β”€ im_eval_callback.py     # Imitation evaluation metrics
β”‚   β”‚   β”œβ”€β”€ im_resample_callback.py # Adaptive motion resampling
β”‚   β”‚   β”œβ”€β”€ model_save_callback.py  # Checkpoint saving
β”‚   β”‚   β”œβ”€β”€ wandb_callback.py       # W&B logging
β”‚   β”‚   └── read_eval_callback.py   # Read eval results from disk
β”‚   └── utils/                  # Math, rotation, scheduling utilities
β”œβ”€β”€ utils/                      # Shared utilities
β”‚   β”œβ”€β”€ motion_lib/             # Motion library loading (PKL format)
β”‚   β”œβ”€β”€ mujoco_sim/             # MuJoCo sim-to-sim bridge
β”‚   └── teleop/                 # VR teleoperation helpers
β”œβ”€β”€ data/                       # Robot models, URDF/USD assets
β”œβ”€β”€ data_process/               # Motion data conversion scripts
└── scripts/                    # MuJoCo sim loop, misc tools
```

---

## Training Pipeline

Running `python gear_sonic/train_agent_trl.py +exp=manager/universal_token/all_modes/sonic_release` executes the following steps:

### 1. Configuration Loading

The entry point uses `@hydra.main(config_path="config", config_name="base")`. The `+exp=...` argument selects an experiment preset that composes all sub-configs:

```
base.yaml                          # Global defaults
  └── +exp=manager/universal_token/all_modes/sonic_release
        β”œβ”€β”€ /algo: ppo_im_phc      # PPO hyperparameters
        β”œβ”€β”€ /actor_critic: universal_token/all_mlp_v1
        β”‚     β”œβ”€β”€ encoders/g1_mf_mlp, smpl_mlp, teleop_mlp
        β”‚     β”œβ”€β”€ decoders/g1_kin_mf_mlp, g1_dyn_mlp
        β”‚     β”œβ”€β”€ quantizers/fsq
        β”‚     └── critics/mlp
        β”œβ”€β”€ /manager_env: base_env  # Environment config
        β”‚     β”œβ”€β”€ observations/{tokenizer, policy, critic}
        β”‚     β”œβ”€β”€ rewards/tracking/base_5point_local_feet_acc
        β”‚     β”œβ”€β”€ terminations/tracking/base_adaptive_strict_ori_foot_xyz
        β”‚     └── events/tracking/level0_4
        β”œβ”€β”€ /aux_losses: universal_token/g1_recon_and_all_latent
        β”œβ”€β”€ /trainer: trl_ppo_aux
        └── /callbacks: model_save, wandb, read_eval, im_resample
```

### 2. Simulator and Accelerator Init

After config resolution, the script:
1. Parses TRL `PPOConfig` / `ScriptArguments` / `ModelConfig` from the config dict.
2. Creates a HuggingFace `Accelerator` for multi-GPU support (DDP).
3. Launches the IsaacLab `AppLauncher` to start the Isaac Sim runtime.
4. Saves `config.yaml` and `meta.yaml` to the experiment directory.

### 3. Environment Creation

`create_manager_env()` instantiates the IsaacLab `ManagerBasedRLEnv` from the composed environment config, then wraps it with `ManagerEnvWrapper`:

```
ManagerBasedRLEnv (IsaacLab)
  └── ManagerEnvWrapper
        β”œβ”€β”€ Observation spaces (policy, critic, tokenizer groups)
        β”œβ”€β”€ Motion command manager (motion_lib)
        β”œβ”€β”€ Action transform module (optional, for pretrained ATM)
        └── Keyboard / visualization hooks
```

### 4. Policy and Value Model Creation

The actor and critic are instantiated from the algo config. For SONIC training, the actor backbone is `UniversalTokenModule`:

```python
# Simplified from train_agent_trl.py
policy = custom_instantiate(config.algo.config.actor, env_config=env.config, ...)
value_model = custom_instantiate(config.algo.config.critic, env_config=env.config, ...)
```

The `Actor` wraps `UniversalTokenModule` as its backbone and adds a diagonal Gaussian distribution for exploration. The `Critic` wraps a separate MLP backbone.

### 5. PPO Training Loop

The `TRLAuxLossPPOTrainer.train()` method runs the main loop:

```
for iteration in range(num_learning_iterations):
    # 1. Rollout: collect num_steps_per_env transitions
    for step in range(num_steps_per_env):
        actions = policy.rollout(obs_dict)
        obs_dict, rewards, dones, infos = env.step(actions)
        store(obs, actions, rewards, values, log_probs)

    # 2. GAE: compute advantages and returns
    advantages = generalized_advantage_estimation(rewards, values, dones)

    # 3. PPO update: num_ppo_epochs over mini-batches
    for epoch in range(num_ppo_epochs):
        for mini_batch in shuffle_and_split(rollout_data):
            policy_loss = clipped_surrogate_objective(...)
            value_loss  = clipped_value_loss(...)
            aux_loss    = sum(coef_i * aux_loss_i)  # encoder reconstruction, etc.
            total_loss  = policy_loss + value_loss_coef * value_loss
                        + aux_loss_scale * aux_loss
            optimizer.step(total_loss)

    # 4. Post-update: sync running stats, adaptive sampling, callbacks
    update_scheduled_params(...)     # learning rate, domain randomization
    callbacks.on_step_end(...)       # checkpointing, evaluation, logging
```

---

## Configuration System

The configuration system uses [Hydra](https://hydra.cc/) with config groups and composition.

### Hierarchy

| Level | Path | Purpose |
|---|---|---|
| **Global** | `config/base.yaml` | Seed, num_envs, paths, wandb toggle |
| **Algorithm** | `config/algo/ppo_im_phc.yaml` | PPO hyperparameters, learning rates, epochs |
| **Actor-Critic** | `config/actor_critic/` | Network architecture (encoders, decoders, critic) |
| **Environment** | `config/manager_env/` | Observations, rewards, terminations, events |
| **Auxiliary Losses** | `config/aux_losses/` | Reconstruction and latent alignment losses |
| **Trainer** | `config/trainer/` | Trainer class selection (PPO or PPO+AuxLoss) |
| **Callbacks** | `config/callbacks/` | Checkpointing, evaluation, W&B logging |
| **Experiment** | `config/exp/` | Preset that composes all the above |

### Experiment Presets

Experiment configs live under `config/exp/` and use the `@package _global_` directive to set values at the root level. They compose all component configs via `defaults`:

```yaml
# config/exp/manager/universal_token/all_modes/sonic_release.yaml
defaults:
  - /algo: ppo_im_phc
  - /manager_env: base_env
  - override /actor_critic: universal_token/all_mlp_v1
  - override /manager_env/observations/tokenizer: unitoken_all_noz
  - override /manager_env/observations/policy: local_dir_hist
  - override /manager_env/rewards: tracking/base_5point_local_feet_acc
  - override /manager_env/terminations: tracking/base_adaptive_strict_ori_foot_xyz
  - override /manager_env/events: tracking/level0_4
  # ...
```

### Key Config Parameters

| Parameter | Default | Description |
|---|---|---|
| `num_envs` | 4096 | Number of parallel simulation environments |
| `algo.config.num_learning_iterations` | 100000 | Total training iterations |
| `algo.config.num_steps_per_env` | 32 | Rollout horizon per iteration |
| `algo.config.num_learning_epochs` | 5 | PPO epochs per iteration |
| `algo.config.num_mini_batches` | 4 | Mini-batches per PPO epoch |
| `algo.config.actor_learning_rate` | 2e-5 | Actor learning rate |
| `algo.config.critic_learning_rate` | 1e-3 | Critic learning rate |
| `algo.config.clip_param` | 0.2 | PPO clipping parameter |
| `algo.config.init_noise_std` | 0.05 | Initial exploration noise std |
| `algo.config.save_interval` | 500 | Checkpoint save frequency (iterations) |

---

## Universal Token Module

The `UniversalTokenModule` implements SONIC's action transform module (ATM) -- the core architecture that maps diverse motion inputs into a shared token space.

### Architecture

```
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  G1 obs    ───►  β”‚  G1 Encoder │──┐
                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  Teleop obs───►  β”‚Teleop Encdr │──┼──► β”‚   FSQ   │──►  β”‚ G1 Dynamic  │──► joint actions
                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚    β”‚Quantizerβ”‚     β”‚   Decoder   β”‚
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  SMPL obs  ───►  β”‚ SMPL Encoderβ”‚β”€β”€β”˜          β”‚
                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜             β”‚         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                              └───────► β”‚G1 Kinematic │──► (aux loss only)
                                                        β”‚   Decoder   β”‚
                                                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```

**Encoders** map different observation modalities into a shared latent space. Each encoder is an MLP that takes modality-specific tokenizer observations and outputs a fixed-size latent vector. During training, one encoder is sampled per environment according to `encoder_sample_probs`.

**FSQ Quantizer** discretizes the continuous latent into a finite set of tokens using Finite Scalar Quantization. Each latent dimension is independently quantized to one of `fsq_level_list` discrete levels. This produces a compact, discrete token representation.

**Decoders** reconstruct outputs from the quantized tokens plus proprioception:
- **G1 Dynamic Decoder** (`g1_dyn`): Produces joint-space actions fed to the actuators. This is the only decoder used at deployment time.
- **G1 Kinematic Decoder** (`g1_kin`): Reconstructs future motion frames from tokens. Used only during training to compute reconstruction auxiliary losses.

### Latent Residual Mode

For downstream tasks (e.g., object manipulation), an external policy can inject corrections into the token space without retraining the base ATM:

| Mode | Behavior |
|---|---|
| `post_quantization` (default) | Residual added after FSQ quantization |
| `pre_quantization` | Residual added before FSQ; the sum gets quantized |
| `pre_quantization_replace` | Latent is replaced entirely by the residual |

### Encoder Sampling

During training, each environment is randomly assigned an encoder per episode according to `encoder_sample_probs`. The `encoder_index` observation tells the module which encoder produced the current token. At deployment, only one encoder is active (selected by the observation configuration).

---

## Environment Structure

The training environment is built on IsaacLab's `ManagerBasedRLEnv` and uses a modular MDP design where each component is configured independently via YAML.

### MDP Components

All MDP components live in `gear_sonic/envs/manager_env/mdp/`:

| Module | Config path | Description |
|---|---|---|
| `observations.py` | `config/manager_env/observations/` | Observation terms for policy, critic, and tokenizer groups |
| `actions.py` | `config/manager_env/actions/` | Joint position action space |
| `rewards.py` | `config/manager_env/rewards/` | Reward terms (tracking, regularization) |
| `terminations.py` | `config/manager_env/terminations/` | Episode termination conditions |
| `events.py` | `config/manager_env/events/` | Domain randomization events |
| `commands.py` | `config/manager_env/commands/` | Motion command generation (motion library) |
| `curriculum.py` | `config/manager_env/curriculum/` | Curriculum schedules |
| `terrain.py` | (inline) | Terrain generation |
| `recorders.py` | `config/manager_env/recorders/` | Video recording |

### Observation Groups

Observations are split into groups, each with its own config file:

| Group | Purpose | Example terms |
|---|---|---|
| **policy** | Direct input to the policy MLP | joint_pos, joint_vel, base_ang_vel, gravity_dir, last_actions |
| **critic** | Privileged observations for the value function | All policy obs + base_lin_vel, body_pos, body_ori |
| **tokenizer** | Input to the UniversalTokenModule encoders | Multi-future joint commands, SMPL joints, VR targets, anchor orientations |

### Reward Terms

Reward configs compose individual terms from `config/manager_env/rewards/terms/`. Key tracking rewards:

| Term | Description |
|---|---|
| `tracking_relative_body_pos` | Track reference body positions (5-point: root, wrists, feet) |
| `tracking_relative_body_ori` | Track reference body orientations |
| `tracking_anchor_pos` | Track root anchor position |
| `tracking_anchor_ori` | Track root anchor orientation |
| `tracking_body_linvel` | Track reference body linear velocities |
| `tracking_body_angvel` | Track reference body angular velocities |
| `action_rate_l2` | Penalize action jerk |
| `feet_acc` | Penalize foot acceleration (smoothness) |

### ManagerEnvWrapper

`ManagerEnvWrapper` bridges the IsaacLab environment with the RL training loop. It handles:
- Flattening observation dicts for the policy
- Applying the optional pretrained action transform module
- Motion replay mode
- Debug visualization and keyboard controls

---

## Evaluation Scripts

### eval_agent_trl.py -- Single Checkpoint

Loads a single checkpoint and runs evaluation in Isaac Sim. Automatically reads the training `config.yaml` from the checkpoint directory to reconstruct the full configuration.

```bash
# Interactive visualization
python gear_sonic/eval_agent_trl.py +checkpoint=path/to/model.pt +headless=False ++num_envs=1

# Headless with video rendering
python gear_sonic/eval_agent_trl.py +checkpoint=path/to/model.pt +headless=True \
    ++num_envs=16 +run_once=True \
    ++manager_env.config.save_rendering_dir=path/to/output \
    ++manager_env.config.render_results=True \
    +manager_env/recorders=render
```

Key features:
- Merges training config with eval overrides (`eval_overrides` in config)
- Removes train-only events and terminations automatically
- Supports `+run_once=True` to exit after all environments complete one episode
- Handles `+metrics_file` to render worst-performing motions from a prior eval

### eval_exp.py -- Checkpoint Monitor

`CheckpointEvaluator` continuously monitors an experiment directory for new checkpoints and evaluates them sequentially. It runs as a companion process alongside training.

```bash
python gear_sonic/eval_exp.py ++experiment_dir=path/to/experiment
```

For each new checkpoint, it:
1. Runs metrics evaluation (launches `eval_agent_trl.py` via subprocess)
2. Runs video rendering for the hardest motions
3. Logs results and videos to W&B (resuming the training run)
4. Marks each checkpoint as evaluated to avoid redundant work

Configuration (`config/eval_exp.yaml`):

| Parameter | Description |
|---|---|
| `experiment_dir` | Path to the training experiment directory |
| `scan_interval` | Seconds between checkpoint scans (default: 60) |
| `num_eval_envs` | Number of environments for metric evaluation |
| `num_render_videos` | Number of videos to render per checkpoint |
| `eval_frequency` | Only evaluate every N-th checkpoint (default: all) |
| `single_pass` | Evaluate pending checkpoints once and exit |

---

## Key Classes Reference

| Class | Module | Description |
|---|---|---|
| `Actor` | `trl/modules/actor_critic_modules.py` | Policy network: backbone + diagonal Gaussian. Maintains observation buffer for temporal models. |
| `Critic` | `trl/modules/actor_critic_modules.py` | Value function network: backbone + scalar output. Supports running mean/std normalization. |
| `UniversalTokenModule` | `trl/modules/universal_token_modules.py` | SONIC ATM: multi-encoder, FSQ quantizer, multi-decoder. Computes auxiliary reconstruction losses. |
| `TRLPPOTrainer` | `trl/trainer/ppo_trainer.py` | Base PPO trainer adapted from HuggingFace TRL. Handles rollout collection, GAE, and gradient updates. |
| `TRLAuxLossPPOTrainer` | `trl/trainer/ppo_trainer_aux_loss.py` | Extends `TRLPPOTrainer` with auxiliary loss support (reconstruction, latent alignment). |
| `PolicyAndValueWrapper` | `trl/trainer/ppo_trainer.py` | Wraps policy + value model into a single `nn.Module` for DDP-safe forward passes. |
| `ManagerEnvWrapper` | `envs/wrapper/manager_env_wrapper.py` | Bridges IsaacLab `ManagerBasedRLEnv` with the training loop. Handles obs flattening, action transforms, replay. |
| `CheckpointEvaluator` | `eval_exp.py` | Monitors experiment directory, evaluates new checkpoints, logs to W&B. |