train_code_for_hs / README.md
dhyun22's picture
README: .single.npz loader alignment + per-view canonical baked in data (no code change)
a392c86 verified
|
Raw History Blame Contribute Delete
12.1 kB
# Training code — CoMind 2-view generation (single-ego setup)
This bundle is our current 2-view egocentric video training stack (Cosmos-Predict2.5, `predict2_multiview`).
It is set up for a **2-view actor-actor** model with a lot of cross-view conditioning. You are going to
**strip it down to a single-ego generation setup that still outputs 2 views**, trained **from scratch on the
CoMind dataset**. This README lists exactly what to change and what to watch out for.
--------------------------------------------------------------------------------
## 0. What the current model conditions on (the parts you will REMOVE / keep)
Per view, each generated clip is conditioned on:
| Signal | Key | What it is | Single-ego plan |
|---|---|---|---|
| warped RGB | `control_input_warped` | RGB warped into the target view from a **combined pool (both views' past)** | **KEEP the channel, but change the DATA**: warp only from **that view's own past context** |
| pose | `control_input_pose` | skeletons. `person_pose=True` = identity-colored, **both people** visible | **KEEP the channel, change the DATA**: render **only the wearer's own** pose |
| in-context references | `reference_frames` (+ `reference_pose`, `reference_cam_w2c`) | clean appearance/pose anchor frames appended as extra tokens | **REMOVE entirely** |
| camera Plücker | `plucker_map` | per-pixel rays. Currently in a **shared cross-view canonical frame (view0 frame0)** | **KEEP** (`enable_plucker=True`), but change the canonical to **each view's OWN first frame** — see §1f |
| reference Plücker | `reference_plucker_map` | posed refs' rays | **REMOVE** (`enable_reference_plucker=False`; no refs) |
| view embedding | net `view_embeddings` | per-view identity added to tokens | **REMOVE** (drop `concat_view_embedding`) |
| cross-view self-attention | token layout `B (V·t) …` | both views attend each other | **Leave as-is** — you do NOT need to separate it; unified attention is fine, and with all cross-view conditioning removed the two views are effectively independent anyway |
So the single-ego model keeps: **warped_cond (own past) + own pose + camera Plücker (own-first-frame canonical)**,
generated per view, no refs / no reference-Plücker / no view-embedding, trained from the Cosmos base 2B.
--------------------------------------------------------------------------------
## 1. Code changes (all in `cosmos_predict2/_src/predict2_multiview/`)
Make a new experiment function in
`configs/vid2vid/experiment/nymeria_pose_2actor.py` (copy `comind_actoractor_personpose_shared` as a start),
and flip these toggles. The flags exist already — you're just turning conditioning OFF.
### 1a. model config (`models/multiview_pose_model_rectified_flow.py` fields, set in the experiment)
```python
cfg["model"]["config"]["num_reference_frames"] = 0 # was 4
cfg["model"]["config"]["enable_reference_plucker"] = False # was True (no refs)
cfg["model"]["config"]["enable_reference_pose"] = False # was True (no refs)
cfg["model"]["config"]["enable_plucker"] = True # KEEP camera Plücker (own-first-frame canonical, §1f)
```
### 1b. net config (`networks/multiview_pose_dit.py`, set via `cfg["model"]["config"]["net"].update(...)`)
```python
cfg["model"]["config"]["net"].update(
enable_reference_frames = False,
num_reference_frames = 0,
shared_reference = False,
enable_reference_plucker= False,
enable_reference_pose = False,
enable_plucker = True, # KEEP Plücker embedder (6-ch ray map -> zero-init add)
concat_view_embedding = False, # <-- drops the view embedding
)
```
`pose_mode` stays `"vae_concat"` (pose is VAE-encoded and concatenated into `cond_embedder`; keep it). The
`cond_embedder` in-channels auto-compute as `warped_latent(16) + visibility(1) + pose_latent(16)` — leave that.
### 1f. Plücker canonical frame = each view's OWN first frame — **already baked in the data, NO code change**
We want each view's Plücker rays expressed relative to its OWN first frame (not a shared cross-view frame).
**The `.single.npz` data already bakes this**: `L_w2c[0]` and `H_w2c[0]` are the identity matrix
(`cam_frame: per_view_canonical_frame0`), i.e. every view's camera poses are already canonicalized to its own
frame0. The model's existing `_preprocess` (~line 185) computes `c2w_ref = inv(w2c[:, 0, 0])`; since view0
frame0 is identity, `c2w_ref = I` and the shared transform is a **no-op**, so each view's rays come out in its
own already-canonical frame — exactly what we want. **Therefore: keep the model Plücker code AS-IS. Do NOT add
a per-view-canonical loop — the data already did it (adding one would be redundant).** Just keep
`enable_plucker=True`, `enable_reference_plucker=False`, and drop the `reference_plucker_map` branch (refs off).
### 1c. checkpoint = from scratch
```python
cfg["checkpoint"]["load_path"] = _base_2b_multiview_ckpt() # Cosmos base 2B, NOT a warm-start
cfg["checkpoint"]["strict_resume"] = False
```
Because refs/Plücker/view-emb are gone and the pose/warp embedders are zero-init or freshly-shaped, **iter-1
loss will start HIGH** (~base video prior) — that's correct for a from-scratch run, unlike a warm-start.
### 1d. data loader
Point `override /data_train` and `override /data_val` at a CoMind data config that:
- reads **your regenerated** CoMind clips (own-past warp + own-only pose),
- sets `num_reference_frames=0`, `shared_reference=False`, `emit_reference_pose=False` (do NOT emit any
`reference_*`), `person_pose=` your choice (you're rendering own-only pose, so the identity-color flag is moot).
The CoMind loader lives in `datasets/comind_pairs.py` / the actor-actor path in `datasets/nymeria_pairs.py`
(`get_nymeria_actor_actor_loader`, class `NymeriaActorActorDataset`). Register a new `cs.store(...)` data
config with the flags above.
### 1e. validation viz
`callbacks/nymeria_validation_viz.py`: set `num_reference_frames=0, shared_reference=False,
emit_reference_pose=False` in the viz kwargs too, or the viz dataset build will look for refs that aren't there.
--------------------------------------------------------------------------------
## 2. DATA — the ablation clips are the sibling `.single.npz` (already being generated)
The single-ego ablation data is stored as a **sibling `<clip>.single.npz`** next to the untouched originals
(`<clip>.npz`, `.campose.npz`, `.refs.npz`). The `.single.npz` is **self-contained** for training and its keys
match the loader's existing convention. Contents (JPEG byte-arrays for image streams, same as the originals):
```
L_/H_target GT video (JPEG) -> video
L_/H_warped OWN-only warp (srcview==0 kept; ==1 -> black hole) -> control_input_warped
L_/H_pose SELF-only skeleton (own hands only) -> control_input_pose
L_/H_vis_packed + vis_shape packed visibility (hole frames vis=False) -> control_input_visibility
L_/H_w2c (77,4,4) PER-VIEW canonical (each view frame0 = identity) -> camera_w2c
K_leader/K_helper (3,3) intrinsics -> camera_K
L_/H_srcf source frame idx (hole=-1) [loader ignores]
warp_mode/pose_mode/cam_frame/n_hole_* markers [loader ignores]
```
**IMPORTANT — source strictly from `.single.npz`.** The original `clip.npz` ALSO has `L_warped/L_pose/
L_vis_packed`, but those are the OLD combined-pool warp + both-people pose. Read warped/pose/vis (and w2c/K)
**only from `.single.npz`**. Since `.single.npz` also carries `L_target`, the simplest wiring is to use
`.single.npz` as the sole source and not open `clip.npz` at all (keypoints aren't read by training). No
`.refs.npz` (references are off).
### 2a. Loader adaptation (small — the format already matches)
Start from `datasets/comind_pairs.py` (`ComindActorActorDataset`), which already reads L_/H_ keys as JPEG
streams. Change:
- **Point the npz path at `<clip>.single.npz`** (and use it for the camera reads too — it has `L_w2c`/
`K_leader`). Drop the `.refs.npz` requirement in the eligibility loop and don't open refs.
- **`person_pose=False`** — `.single.npz` has `L_pose` (self-only), NOT `L_pose_person`; with
`person_pose=True` the loader would KeyError.
- **`num_reference_frames=0`** (R=0) so the reference block is skipped.
- Everything else is unchanged: `_stack_jpeg` decode, `np.unpackbits(vis_packed)` with `vis_shape=(77,504,504)`,
`_fit_poses(w2c,(77))`, `_resize` to your train resolution — all already match the `.single.npz` layout.
Verified against a real file (`62531cc1_006159.single.npz`, 28.5MB): warped/pose/target are JPEG object arrays;
`vis_shape=(77,504,504)` with packed-bit capacity exactly `77*504*504`; `L_w2c[0]==H_w2c[0]==identity`
(per-view canonical). So with the three config flips above, the existing loader/decoder path lines up 1:1.
--------------------------------------------------------------------------------
## 3. Things to know / gotchas
- **Env / launch**: `sh/train_nymeria_longer.sh`, run as
`CUDA_VISIBLE_DEVICES=0,1,2,3 EXP=<your_exp_name> NPROC=4 bash sh/train_nymeria_longer.sh`.
Conda env `ego_dh`. Needs `COSMOS_QWEN_TOKENIZER_DIR=/data/cosmos_reason1_7b` for the online text encoder,
`IMAGINAIRE_OUTPUT_ROOT=<out>`.
- **Online text encoding**: CoMind emits `ai_caption`; the model text-encodes it online each step (Qwen). Keep
the tokenizer dir env set or conditioning will crash on `t5_text_embeddings=None`.
- **DCP checkpoints** are reshardable across NPROC (4↔3↔2 resume works). `save_iter` in the experiment.
- **register the experiment**: append your function to the `experiments = [...]` list at the bottom of
`nymeria_pose_2actor.py`, and confirm it builds:
`python -c "..."` overriding `experiment=<name>` (see how the file builds configs) before launching 4-GPU.
- **`state_t`**: latent temporal length = `1 + (T-1)//4`. For 77 frames it's 20. It appears in the model config;
keep it consistent between data (`num_video_frames`) and model.
- **AR / noise recipe**: `predict2/models/video2world_model_rectified_flow.py` has an optional
`noisy_conditioning_*` + `conditional_frames_probs` recipe (teacher-forcing for autoregressive rollout). It is
OFF by default (`{1:1.0}`, prob 0). Ignore it unless you want AR — not needed for the single-ego baseline.
- **Cross-attention**: you asked whether to separate it — you don't need to. The DiT already runs unified
self-attention over `B (V·t) H W D`; with no cross-view conditioning the two views just don't share useful
signal, which is exactly the single-ego behavior. Leaving it unified is simplest and correct.
--------------------------------------------------------------------------------
## 4. File map (what's in this bundle)
```
cosmos_predict2/_src/predict2_multiview/
configs/vid2vid/experiment/nymeria_pose_2actor.py # all experiments (copy comind_* -> your single-ego exp)
configs/vid2vid/defaults/{conditioner,data,model}.py
datasets/nymeria_pairs.py # NymeriaActorActorDataset + loaders + cs.store data configs
datasets/comind_pairs.py # CoMind dataset/loader
networks/multiview_pose_dit.py # the DiT: all the enable_* / view-emb / reference embedders
models/multiview_pose_model_rectified_flow.py # conditioning preprocessing (refs/plucker/refpose/depth)
models/multiview_vid2vid_model_rectified_flow.py
callbacks/nymeria_validation_viz.py # validation grid
predict2/models/video2world_model_rectified_flow.py # base rectified-flow denoise (+optional AR noise)
sh/train_nymeria_longer.sh # launch
```
Summary of the diff you need: **turn OFF** `enable_reference_frames`, `enable_reference_plucker`,
`enable_reference_pose`, `num_reference_frames=0`, `shared_reference=False`, `concat_view_embedding=False`;
**KEEP `enable_plucker=True` but switch its canonical to each view's own first frame (§1f)**; **from base 2B**;
and feed **CoMind clips rebuilt with own-past warp + own-only pose**. Keep warped_cond + pose + Plücker
channels and unified cross-view attention.