# Training code — CoMind 2-view generation (single-ego setup) This bundle is our current 2-view egocentric video training stack (Cosmos-Predict2.5, `predict2_multiview`). It is set up for a **2-view actor-actor** model with a lot of cross-view conditioning. You are going to **strip it down to a single-ego generation setup that still outputs 2 views**, trained **from scratch on the CoMind dataset**. This README lists exactly what to change and what to watch out for. -------------------------------------------------------------------------------- ## 0. What the current model conditions on (the parts you will REMOVE / keep) Per view, each generated clip is conditioned on: | Signal | Key | What it is | Single-ego plan | |---|---|---|---| | warped RGB | `control_input_warped` | RGB warped into the target view from a **combined pool (both views' past)** | **KEEP the channel, but change the DATA**: warp only from **that view's own past context** | | pose | `control_input_pose` | skeletons. `person_pose=True` = identity-colored, **both people** visible | **KEEP the channel, change the DATA**: render **only the wearer's own** pose | | in-context references | `reference_frames` (+ `reference_pose`, `reference_cam_w2c`) | clean appearance/pose anchor frames appended as extra tokens | **REMOVE entirely** | | camera Plücker | `plucker_map` | per-pixel rays. Currently in a **shared cross-view canonical frame (view0 frame0)** | **KEEP** (`enable_plucker=True`), but change the canonical to **each view's OWN first frame** — see §1f | | reference Plücker | `reference_plucker_map` | posed refs' rays | **REMOVE** (`enable_reference_plucker=False`; no refs) | | view embedding | net `view_embeddings` | per-view identity added to tokens | **REMOVE** (drop `concat_view_embedding`) | | cross-view self-attention | token layout `B (V·t) …` | both views attend each other | **Leave as-is** — you do NOT need to separate it; unified attention is fine, and with all cross-view conditioning removed the two views are effectively independent anyway | So the single-ego model keeps: **warped_cond (own past) + own pose + camera Plücker (own-first-frame canonical)**, generated per view, no refs / no reference-Plücker / no view-embedding, trained from the Cosmos base 2B. -------------------------------------------------------------------------------- ## 1. Code changes (all in `cosmos_predict2/_src/predict2_multiview/`) Make a new experiment function in `configs/vid2vid/experiment/nymeria_pose_2actor.py` (copy `comind_actoractor_personpose_shared` as a start), and flip these toggles. The flags exist already — you're just turning conditioning OFF. ### 1a. model config (`models/multiview_pose_model_rectified_flow.py` fields, set in the experiment) ```python cfg["model"]["config"]["num_reference_frames"] = 0 # was 4 cfg["model"]["config"]["enable_reference_plucker"] = False # was True (no refs) cfg["model"]["config"]["enable_reference_pose"] = False # was True (no refs) cfg["model"]["config"]["enable_plucker"] = True # KEEP camera Plücker (own-first-frame canonical, §1f) ``` ### 1b. net config (`networks/multiview_pose_dit.py`, set via `cfg["model"]["config"]["net"].update(...)`) ```python cfg["model"]["config"]["net"].update( enable_reference_frames = False, num_reference_frames = 0, shared_reference = False, enable_reference_plucker= False, enable_reference_pose = False, enable_plucker = True, # KEEP Plücker embedder (6-ch ray map -> zero-init add) concat_view_embedding = False, # <-- drops the view embedding ) ``` `pose_mode` stays `"vae_concat"` (pose is VAE-encoded and concatenated into `cond_embedder`; keep it). The `cond_embedder` in-channels auto-compute as `warped_latent(16) + visibility(1) + pose_latent(16)` — leave that. ### 1f. Plücker canonical frame = each view's OWN first frame — **already baked in the data, NO code change** We want each view's Plücker rays expressed relative to its OWN first frame (not a shared cross-view frame). **The `.single.npz` data already bakes this**: `L_w2c[0]` and `H_w2c[0]` are the identity matrix (`cam_frame: per_view_canonical_frame0`), i.e. every view's camera poses are already canonicalized to its own frame0. The model's existing `_preprocess` (~line 185) computes `c2w_ref = inv(w2c[:, 0, 0])`; since view0 frame0 is identity, `c2w_ref = I` and the shared transform is a **no-op**, so each view's rays come out in its own already-canonical frame — exactly what we want. **Therefore: keep the model Plücker code AS-IS. Do NOT add a per-view-canonical loop — the data already did it (adding one would be redundant).** Just keep `enable_plucker=True`, `enable_reference_plucker=False`, and drop the `reference_plucker_map` branch (refs off). ### 1c. checkpoint = from scratch ```python cfg["checkpoint"]["load_path"] = _base_2b_multiview_ckpt() # Cosmos base 2B, NOT a warm-start cfg["checkpoint"]["strict_resume"] = False ``` Because refs/Plücker/view-emb are gone and the pose/warp embedders are zero-init or freshly-shaped, **iter-1 loss will start HIGH** (~base video prior) — that's correct for a from-scratch run, unlike a warm-start. ### 1d. data loader Point `override /data_train` and `override /data_val` at a CoMind data config that: - reads **your regenerated** CoMind clips (own-past warp + own-only pose), - sets `num_reference_frames=0`, `shared_reference=False`, `emit_reference_pose=False` (do NOT emit any `reference_*`), `person_pose=` your choice (you're rendering own-only pose, so the identity-color flag is moot). The CoMind loader lives in `datasets/comind_pairs.py` / the actor-actor path in `datasets/nymeria_pairs.py` (`get_nymeria_actor_actor_loader`, class `NymeriaActorActorDataset`). Register a new `cs.store(...)` data config with the flags above. ### 1e. validation viz `callbacks/nymeria_validation_viz.py`: set `num_reference_frames=0, shared_reference=False, emit_reference_pose=False` in the viz kwargs too, or the viz dataset build will look for refs that aren't there. -------------------------------------------------------------------------------- ## 2. DATA — the ablation clips are the sibling `.single.npz` (already being generated) The single-ego ablation data is stored as a **sibling `.single.npz`** next to the untouched originals (`.npz`, `.campose.npz`, `.refs.npz`). The `.single.npz` is **self-contained** for training and its keys match the loader's existing convention. Contents (JPEG byte-arrays for image streams, same as the originals): ``` L_/H_target GT video (JPEG) -> video L_/H_warped OWN-only warp (srcview==0 kept; ==1 -> black hole) -> control_input_warped L_/H_pose SELF-only skeleton (own hands only) -> control_input_pose L_/H_vis_packed + vis_shape packed visibility (hole frames vis=False) -> control_input_visibility L_/H_w2c (77,4,4) PER-VIEW canonical (each view frame0 = identity) -> camera_w2c K_leader/K_helper (3,3) intrinsics -> camera_K L_/H_srcf source frame idx (hole=-1) [loader ignores] warp_mode/pose_mode/cam_frame/n_hole_* markers [loader ignores] ``` **IMPORTANT — source strictly from `.single.npz`.** The original `clip.npz` ALSO has `L_warped/L_pose/ L_vis_packed`, but those are the OLD combined-pool warp + both-people pose. Read warped/pose/vis (and w2c/K) **only from `.single.npz`**. Since `.single.npz` also carries `L_target`, the simplest wiring is to use `.single.npz` as the sole source and not open `clip.npz` at all (keypoints aren't read by training). No `.refs.npz` (references are off). ### 2a. Loader adaptation (small — the format already matches) Start from `datasets/comind_pairs.py` (`ComindActorActorDataset`), which already reads L_/H_ keys as JPEG streams. Change: - **Point the npz path at `.single.npz`** (and use it for the camera reads too — it has `L_w2c`/ `K_leader`). Drop the `.refs.npz` requirement in the eligibility loop and don't open refs. - **`person_pose=False`** — `.single.npz` has `L_pose` (self-only), NOT `L_pose_person`; with `person_pose=True` the loader would KeyError. - **`num_reference_frames=0`** (R=0) so the reference block is skipped. - Everything else is unchanged: `_stack_jpeg` decode, `np.unpackbits(vis_packed)` with `vis_shape=(77,504,504)`, `_fit_poses(w2c,(77))`, `_resize` to your train resolution — all already match the `.single.npz` layout. Verified against a real file (`62531cc1_006159.single.npz`, 28.5MB): warped/pose/target are JPEG object arrays; `vis_shape=(77,504,504)` with packed-bit capacity exactly `77*504*504`; `L_w2c[0]==H_w2c[0]==identity` (per-view canonical). So with the three config flips above, the existing loader/decoder path lines up 1:1. -------------------------------------------------------------------------------- ## 3. Things to know / gotchas - **Env / launch**: `sh/train_nymeria_longer.sh`, run as `CUDA_VISIBLE_DEVICES=0,1,2,3 EXP= NPROC=4 bash sh/train_nymeria_longer.sh`. Conda env `ego_dh`. Needs `COSMOS_QWEN_TOKENIZER_DIR=/data/cosmos_reason1_7b` for the online text encoder, `IMAGINAIRE_OUTPUT_ROOT=`. - **Online text encoding**: CoMind emits `ai_caption`; the model text-encodes it online each step (Qwen). Keep the tokenizer dir env set or conditioning will crash on `t5_text_embeddings=None`. - **DCP checkpoints** are reshardable across NPROC (4↔3↔2 resume works). `save_iter` in the experiment. - **register the experiment**: append your function to the `experiments = [...]` list at the bottom of `nymeria_pose_2actor.py`, and confirm it builds: `python -c "..."` overriding `experiment=` (see how the file builds configs) before launching 4-GPU. - **`state_t`**: latent temporal length = `1 + (T-1)//4`. For 77 frames it's 20. It appears in the model config; keep it consistent between data (`num_video_frames`) and model. - **AR / noise recipe**: `predict2/models/video2world_model_rectified_flow.py` has an optional `noisy_conditioning_*` + `conditional_frames_probs` recipe (teacher-forcing for autoregressive rollout). It is OFF by default (`{1:1.0}`, prob 0). Ignore it unless you want AR — not needed for the single-ego baseline. - **Cross-attention**: you asked whether to separate it — you don't need to. The DiT already runs unified self-attention over `B (V·t) H W D`; with no cross-view conditioning the two views just don't share useful signal, which is exactly the single-ego behavior. Leaving it unified is simplest and correct. -------------------------------------------------------------------------------- ## 4. File map (what's in this bundle) ``` cosmos_predict2/_src/predict2_multiview/ configs/vid2vid/experiment/nymeria_pose_2actor.py # all experiments (copy comind_* -> your single-ego exp) configs/vid2vid/defaults/{conditioner,data,model}.py datasets/nymeria_pairs.py # NymeriaActorActorDataset + loaders + cs.store data configs datasets/comind_pairs.py # CoMind dataset/loader networks/multiview_pose_dit.py # the DiT: all the enable_* / view-emb / reference embedders models/multiview_pose_model_rectified_flow.py # conditioning preprocessing (refs/plucker/refpose/depth) models/multiview_vid2vid_model_rectified_flow.py callbacks/nymeria_validation_viz.py # validation grid predict2/models/video2world_model_rectified_flow.py # base rectified-flow denoise (+optional AR noise) sh/train_nymeria_longer.sh # launch ``` Summary of the diff you need: **turn OFF** `enable_reference_frames`, `enable_reference_plucker`, `enable_reference_pose`, `num_reference_frames=0`, `shared_reference=False`, `concat_view_embedding=False`; **KEEP `enable_plucker=True` but switch its canonical to each view's own first frame (§1f)**; **from base 2B**; and feed **CoMind clips rebuilt with own-past warp + own-only pose**. Keep warped_cond + pose + Plücker channels and unified cross-view attention.