|
Download README.md from dhyun22/train_code_for_hs: direct link, hf CLI and curl.
- Browser
- Download file 12.1 kB
-
https://huggingface.co/dhyun22/train_code_for_hs/resolve/main/README.md
- Command line
-
hf download hf://dhyun22/train_code_for_hs/README.md
-
curl -L -o README.md https://huggingface.co/dhyun22/train_code_for_hs/resolve/main/README.md
12.1 kB
| # Training code — CoMind 2-view generation (single-ego setup) | |
| This bundle is our current 2-view egocentric video training stack (Cosmos-Predict2.5, `predict2_multiview`). | |
| It is set up for a **2-view actor-actor** model with a lot of cross-view conditioning. You are going to | |
| **strip it down to a single-ego generation setup that still outputs 2 views**, trained **from scratch on the | |
| CoMind dataset**. This README lists exactly what to change and what to watch out for. | |
| -------------------------------------------------------------------------------- | |
| ## 0. What the current model conditions on (the parts you will REMOVE / keep) | |
| Per view, each generated clip is conditioned on: | |
| | Signal | Key | What it is | Single-ego plan | | |
| |---|---|---|---| | |
| | warped RGB | `control_input_warped` | RGB warped into the target view from a **combined pool (both views' past)** | **KEEP the channel, but change the DATA**: warp only from **that view's own past context** | | |
| | pose | `control_input_pose` | skeletons. `person_pose=True` = identity-colored, **both people** visible | **KEEP the channel, change the DATA**: render **only the wearer's own** pose | | |
| | in-context references | `reference_frames` (+ `reference_pose`, `reference_cam_w2c`) | clean appearance/pose anchor frames appended as extra tokens | **REMOVE entirely** | | |
| | camera Plücker | `plucker_map` | per-pixel rays. Currently in a **shared cross-view canonical frame (view0 frame0)** | **KEEP** (`enable_plucker=True`), but change the canonical to **each view's OWN first frame** — see §1f | | |
| | reference Plücker | `reference_plucker_map` | posed refs' rays | **REMOVE** (`enable_reference_plucker=False`; no refs) | | |
| | view embedding | net `view_embeddings` | per-view identity added to tokens | **REMOVE** (drop `concat_view_embedding`) | | |
| | cross-view self-attention | token layout `B (V·t) …` | both views attend each other | **Leave as-is** — you do NOT need to separate it; unified attention is fine, and with all cross-view conditioning removed the two views are effectively independent anyway | | |
| So the single-ego model keeps: **warped_cond (own past) + own pose + camera Plücker (own-first-frame canonical)**, | |
| generated per view, no refs / no reference-Plücker / no view-embedding, trained from the Cosmos base 2B. | |
| -------------------------------------------------------------------------------- | |
| ## 1. Code changes (all in `cosmos_predict2/_src/predict2_multiview/`) | |
| Make a new experiment function in | |
| `configs/vid2vid/experiment/nymeria_pose_2actor.py` (copy `comind_actoractor_personpose_shared` as a start), | |
| and flip these toggles. The flags exist already — you're just turning conditioning OFF. | |
| ### 1a. model config (`models/multiview_pose_model_rectified_flow.py` fields, set in the experiment) | |
| ```python | |
| cfg["model"]["config"]["num_reference_frames"] = 0 # was 4 | |
| cfg["model"]["config"]["enable_reference_plucker"] = False # was True (no refs) | |
| cfg["model"]["config"]["enable_reference_pose"] = False # was True (no refs) | |
| cfg["model"]["config"]["enable_plucker"] = True # KEEP camera Plücker (own-first-frame canonical, §1f) | |
| ``` | |
| ### 1b. net config (`networks/multiview_pose_dit.py`, set via `cfg["model"]["config"]["net"].update(...)`) | |
| ```python | |
| cfg["model"]["config"]["net"].update( | |
| enable_reference_frames = False, | |
| num_reference_frames = 0, | |
| shared_reference = False, | |
| enable_reference_plucker= False, | |
| enable_reference_pose = False, | |
| enable_plucker = True, # KEEP Plücker embedder (6-ch ray map -> zero-init add) | |
| concat_view_embedding = False, # <-- drops the view embedding | |
| ) | |
| ``` | |
| `pose_mode` stays `"vae_concat"` (pose is VAE-encoded and concatenated into `cond_embedder`; keep it). The | |
| `cond_embedder` in-channels auto-compute as `warped_latent(16) + visibility(1) + pose_latent(16)` — leave that. | |
| ### 1f. Plücker canonical frame = each view's OWN first frame — **already baked in the data, NO code change** | |
| We want each view's Plücker rays expressed relative to its OWN first frame (not a shared cross-view frame). | |
| **The `.single.npz` data already bakes this**: `L_w2c[0]` and `H_w2c[0]` are the identity matrix | |
| (`cam_frame: per_view_canonical_frame0`), i.e. every view's camera poses are already canonicalized to its own | |
| frame0. The model's existing `_preprocess` (~line 185) computes `c2w_ref = inv(w2c[:, 0, 0])`; since view0 | |
| frame0 is identity, `c2w_ref = I` and the shared transform is a **no-op**, so each view's rays come out in its | |
| own already-canonical frame — exactly what we want. **Therefore: keep the model Plücker code AS-IS. Do NOT add | |
| a per-view-canonical loop — the data already did it (adding one would be redundant).** Just keep | |
| `enable_plucker=True`, `enable_reference_plucker=False`, and drop the `reference_plucker_map` branch (refs off). | |
| ### 1c. checkpoint = from scratch | |
| ```python | |
| cfg["checkpoint"]["load_path"] = _base_2b_multiview_ckpt() # Cosmos base 2B, NOT a warm-start | |
| cfg["checkpoint"]["strict_resume"] = False | |
| ``` | |
| Because refs/Plücker/view-emb are gone and the pose/warp embedders are zero-init or freshly-shaped, **iter-1 | |
| loss will start HIGH** (~base video prior) — that's correct for a from-scratch run, unlike a warm-start. | |
| ### 1d. data loader | |
| Point `override /data_train` and `override /data_val` at a CoMind data config that: | |
| - reads **your regenerated** CoMind clips (own-past warp + own-only pose), | |
| - sets `num_reference_frames=0`, `shared_reference=False`, `emit_reference_pose=False` (do NOT emit any | |
| `reference_*`), `person_pose=` your choice (you're rendering own-only pose, so the identity-color flag is moot). | |
| The CoMind loader lives in `datasets/comind_pairs.py` / the actor-actor path in `datasets/nymeria_pairs.py` | |
| (`get_nymeria_actor_actor_loader`, class `NymeriaActorActorDataset`). Register a new `cs.store(...)` data | |
| config with the flags above. | |
| ### 1e. validation viz | |
| `callbacks/nymeria_validation_viz.py`: set `num_reference_frames=0, shared_reference=False, | |
| emit_reference_pose=False` in the viz kwargs too, or the viz dataset build will look for refs that aren't there. | |
| -------------------------------------------------------------------------------- | |
| ## 2. DATA — the ablation clips are the sibling `.single.npz` (already being generated) | |
| The single-ego ablation data is stored as a **sibling `<clip>.single.npz`** next to the untouched originals | |
| (`<clip>.npz`, `.campose.npz`, `.refs.npz`). The `.single.npz` is **self-contained** for training and its keys | |
| match the loader's existing convention. Contents (JPEG byte-arrays for image streams, same as the originals): | |
| ``` | |
| L_/H_target GT video (JPEG) -> video | |
| L_/H_warped OWN-only warp (srcview==0 kept; ==1 -> black hole) -> control_input_warped | |
| L_/H_pose SELF-only skeleton (own hands only) -> control_input_pose | |
| L_/H_vis_packed + vis_shape packed visibility (hole frames vis=False) -> control_input_visibility | |
| L_/H_w2c (77,4,4) PER-VIEW canonical (each view frame0 = identity) -> camera_w2c | |
| K_leader/K_helper (3,3) intrinsics -> camera_K | |
| L_/H_srcf source frame idx (hole=-1) [loader ignores] | |
| warp_mode/pose_mode/cam_frame/n_hole_* markers [loader ignores] | |
| ``` | |
| **IMPORTANT — source strictly from `.single.npz`.** The original `clip.npz` ALSO has `L_warped/L_pose/ | |
| L_vis_packed`, but those are the OLD combined-pool warp + both-people pose. Read warped/pose/vis (and w2c/K) | |
| **only from `.single.npz`**. Since `.single.npz` also carries `L_target`, the simplest wiring is to use | |
| `.single.npz` as the sole source and not open `clip.npz` at all (keypoints aren't read by training). No | |
| `.refs.npz` (references are off). | |
| ### 2a. Loader adaptation (small — the format already matches) | |
| Start from `datasets/comind_pairs.py` (`ComindActorActorDataset`), which already reads L_/H_ keys as JPEG | |
| streams. Change: | |
| - **Point the npz path at `<clip>.single.npz`** (and use it for the camera reads too — it has `L_w2c`/ | |
| `K_leader`). Drop the `.refs.npz` requirement in the eligibility loop and don't open refs. | |
| - **`person_pose=False`** — `.single.npz` has `L_pose` (self-only), NOT `L_pose_person`; with | |
| `person_pose=True` the loader would KeyError. | |
| - **`num_reference_frames=0`** (R=0) so the reference block is skipped. | |
| - Everything else is unchanged: `_stack_jpeg` decode, `np.unpackbits(vis_packed)` with `vis_shape=(77,504,504)`, | |
| `_fit_poses(w2c,(77))`, `_resize` to your train resolution — all already match the `.single.npz` layout. | |
| Verified against a real file (`62531cc1_006159.single.npz`, 28.5MB): warped/pose/target are JPEG object arrays; | |
| `vis_shape=(77,504,504)` with packed-bit capacity exactly `77*504*504`; `L_w2c[0]==H_w2c[0]==identity` | |
| (per-view canonical). So with the three config flips above, the existing loader/decoder path lines up 1:1. | |
| -------------------------------------------------------------------------------- | |
| ## 3. Things to know / gotchas | |
| - **Env / launch**: `sh/train_nymeria_longer.sh`, run as | |
| `CUDA_VISIBLE_DEVICES=0,1,2,3 EXP=<your_exp_name> NPROC=4 bash sh/train_nymeria_longer.sh`. | |
| Conda env `ego_dh`. Needs `COSMOS_QWEN_TOKENIZER_DIR=/data/cosmos_reason1_7b` for the online text encoder, | |
| `IMAGINAIRE_OUTPUT_ROOT=<out>`. | |
| - **Online text encoding**: CoMind emits `ai_caption`; the model text-encodes it online each step (Qwen). Keep | |
| the tokenizer dir env set or conditioning will crash on `t5_text_embeddings=None`. | |
| - **DCP checkpoints** are reshardable across NPROC (4↔3↔2 resume works). `save_iter` in the experiment. | |
| - **register the experiment**: append your function to the `experiments = [...]` list at the bottom of | |
| `nymeria_pose_2actor.py`, and confirm it builds: | |
| `python -c "..."` overriding `experiment=<name>` (see how the file builds configs) before launching 4-GPU. | |
| - **`state_t`**: latent temporal length = `1 + (T-1)//4`. For 77 frames it's 20. It appears in the model config; | |
| keep it consistent between data (`num_video_frames`) and model. | |
| - **AR / noise recipe**: `predict2/models/video2world_model_rectified_flow.py` has an optional | |
| `noisy_conditioning_*` + `conditional_frames_probs` recipe (teacher-forcing for autoregressive rollout). It is | |
| OFF by default (`{1:1.0}`, prob 0). Ignore it unless you want AR — not needed for the single-ego baseline. | |
| - **Cross-attention**: you asked whether to separate it — you don't need to. The DiT already runs unified | |
| self-attention over `B (V·t) H W D`; with no cross-view conditioning the two views just don't share useful | |
| signal, which is exactly the single-ego behavior. Leaving it unified is simplest and correct. | |
| -------------------------------------------------------------------------------- | |
| ## 4. File map (what's in this bundle) | |
| ``` | |
| cosmos_predict2/_src/predict2_multiview/ | |
| configs/vid2vid/experiment/nymeria_pose_2actor.py # all experiments (copy comind_* -> your single-ego exp) | |
| configs/vid2vid/defaults/{conditioner,data,model}.py | |
| datasets/nymeria_pairs.py # NymeriaActorActorDataset + loaders + cs.store data configs | |
| datasets/comind_pairs.py # CoMind dataset/loader | |
| networks/multiview_pose_dit.py # the DiT: all the enable_* / view-emb / reference embedders | |
| models/multiview_pose_model_rectified_flow.py # conditioning preprocessing (refs/plucker/refpose/depth) | |
| models/multiview_vid2vid_model_rectified_flow.py | |
| callbacks/nymeria_validation_viz.py # validation grid | |
| predict2/models/video2world_model_rectified_flow.py # base rectified-flow denoise (+optional AR noise) | |
| sh/train_nymeria_longer.sh # launch | |
| ``` | |
| Summary of the diff you need: **turn OFF** `enable_reference_frames`, `enable_reference_plucker`, | |
| `enable_reference_pose`, `num_reference_frames=0`, `shared_reference=False`, `concat_view_embedding=False`; | |
| **KEEP `enable_plucker=True` but switch its canonical to each view's own first frame (§1f)**; **from base 2B**; | |
| and feed **CoMind clips rebuilt with own-past warp + own-only pose**. Keep warped_cond + pose + Plücker | |
| channels and unified cross-view attention. | |