train_code_for_hs / README.md
dhyun22's picture
README: .single.npz loader alignment + per-view canonical baked in data (no code change)
a392c86 verified
|
Raw History Blame Contribute Delete
12.1 kB

Training code — CoMind 2-view generation (single-ego setup)

This bundle is our current 2-view egocentric video training stack (Cosmos-Predict2.5, predict2_multiview). It is set up for a 2-view actor-actor model with a lot of cross-view conditioning. You are going to strip it down to a single-ego generation setup that still outputs 2 views, trained from scratch on the CoMind dataset. This README lists exactly what to change and what to watch out for.


0. What the current model conditions on (the parts you will REMOVE / keep)

Per view, each generated clip is conditioned on:

Signal Key What it is Single-ego plan
warped RGB control_input_warped RGB warped into the target view from a combined pool (both views' past) KEEP the channel, but change the DATA: warp only from that view's own past context
pose control_input_pose skeletons. person_pose=True = identity-colored, both people visible KEEP the channel, change the DATA: render only the wearer's own pose
in-context references reference_frames (+ reference_pose, reference_cam_w2c) clean appearance/pose anchor frames appended as extra tokens REMOVE entirely
camera Plücker plucker_map per-pixel rays. Currently in a shared cross-view canonical frame (view0 frame0) KEEP (enable_plucker=True), but change the canonical to each view's OWN first frame — see §1f
reference Plücker reference_plucker_map posed refs' rays REMOVE (enable_reference_plucker=False; no refs)
view embedding net view_embeddings per-view identity added to tokens REMOVE (drop concat_view_embedding)
cross-view self-attention token layout B (V·t) … both views attend each other Leave as-is — you do NOT need to separate it; unified attention is fine, and with all cross-view conditioning removed the two views are effectively independent anyway

So the single-ego model keeps: warped_cond (own past) + own pose + camera Plücker (own-first-frame canonical), generated per view, no refs / no reference-Plücker / no view-embedding, trained from the Cosmos base 2B.


1. Code changes (all in cosmos_predict2/_src/predict2_multiview/)

Make a new experiment function in configs/vid2vid/experiment/nymeria_pose_2actor.py (copy comind_actoractor_personpose_shared as a start), and flip these toggles. The flags exist already — you're just turning conditioning OFF.

1a. model config (models/multiview_pose_model_rectified_flow.py fields, set in the experiment)

cfg["model"]["config"]["num_reference_frames"]      = 0       # was 4
cfg["model"]["config"]["enable_reference_plucker"]  = False   # was True (no refs)
cfg["model"]["config"]["enable_reference_pose"]     = False   # was True (no refs)
cfg["model"]["config"]["enable_plucker"]            = True    # KEEP camera Plücker (own-first-frame canonical, §1f)

1b. net config (networks/multiview_pose_dit.py, set via cfg["model"]["config"]["net"].update(...))

cfg["model"]["config"]["net"].update(
    enable_reference_frames = False,
    num_reference_frames    = 0,
    shared_reference        = False,
    enable_reference_plucker= False,
    enable_reference_pose   = False,
    enable_plucker          = True,    # KEEP Plücker embedder (6-ch ray map -> zero-init add)
    concat_view_embedding   = False,   # <-- drops the view embedding
)

pose_mode stays "vae_concat" (pose is VAE-encoded and concatenated into cond_embedder; keep it). The cond_embedder in-channels auto-compute as warped_latent(16) + visibility(1) + pose_latent(16) — leave that.

1f. Plücker canonical frame = each view's OWN first frame — already baked in the data, NO code change

We want each view's Plücker rays expressed relative to its OWN first frame (not a shared cross-view frame). The .single.npz data already bakes this: L_w2c[0] and H_w2c[0] are the identity matrix (cam_frame: per_view_canonical_frame0), i.e. every view's camera poses are already canonicalized to its own frame0. The model's existing _preprocess (~line 185) computes c2w_ref = inv(w2c[:, 0, 0]); since view0 frame0 is identity, c2w_ref = I and the shared transform is a no-op, so each view's rays come out in its own already-canonical frame — exactly what we want. Therefore: keep the model Plücker code AS-IS. Do NOT add a per-view-canonical loop — the data already did it (adding one would be redundant). Just keep enable_plucker=True, enable_reference_plucker=False, and drop the reference_plucker_map branch (refs off).

1c. checkpoint = from scratch

cfg["checkpoint"]["load_path"]      = _base_2b_multiview_ckpt()   # Cosmos base 2B, NOT a warm-start
cfg["checkpoint"]["strict_resume"]  = False

Because refs/Plücker/view-emb are gone and the pose/warp embedders are zero-init or freshly-shaped, iter-1 loss will start HIGH (~base video prior) — that's correct for a from-scratch run, unlike a warm-start.

1d. data loader

Point override /data_train and override /data_val at a CoMind data config that:

  • reads your regenerated CoMind clips (own-past warp + own-only pose),
  • sets num_reference_frames=0, shared_reference=False, emit_reference_pose=False (do NOT emit any reference_*), person_pose= your choice (you're rendering own-only pose, so the identity-color flag is moot).

The CoMind loader lives in datasets/comind_pairs.py / the actor-actor path in datasets/nymeria_pairs.py (get_nymeria_actor_actor_loader, class NymeriaActorActorDataset). Register a new cs.store(...) data config with the flags above.

1e. validation viz

callbacks/nymeria_validation_viz.py: set num_reference_frames=0, shared_reference=False, emit_reference_pose=False in the viz kwargs too, or the viz dataset build will look for refs that aren't there.


2. DATA — the ablation clips are the sibling .single.npz (already being generated)

The single-ego ablation data is stored as a sibling <clip>.single.npz next to the untouched originals (<clip>.npz, .campose.npz, .refs.npz). The .single.npz is self-contained for training and its keys match the loader's existing convention. Contents (JPEG byte-arrays for image streams, same as the originals):

L_/H_target        GT video (JPEG)                         -> video
L_/H_warped        OWN-only warp (srcview==0 kept; ==1 -> black hole)  -> control_input_warped
L_/H_pose          SELF-only skeleton (own hands only)      -> control_input_pose
L_/H_vis_packed + vis_shape   packed visibility (hole frames vis=False) -> control_input_visibility
L_/H_w2c (77,4,4)  PER-VIEW canonical (each view frame0 = identity)     -> camera_w2c
K_leader/K_helper (3,3)   intrinsics                        -> camera_K
L_/H_srcf          source frame idx (hole=-1)  [loader ignores]
warp_mode/pose_mode/cam_frame/n_hole_*   markers [loader ignores]

IMPORTANT — source strictly from .single.npz. The original clip.npz ALSO has L_warped/L_pose/ L_vis_packed, but those are the OLD combined-pool warp + both-people pose. Read warped/pose/vis (and w2c/K) only from .single.npz. Since .single.npz also carries L_target, the simplest wiring is to use .single.npz as the sole source and not open clip.npz at all (keypoints aren't read by training). No .refs.npz (references are off).

2a. Loader adaptation (small — the format already matches)

Start from datasets/comind_pairs.py (ComindActorActorDataset), which already reads L_/H_ keys as JPEG streams. Change:

  • Point the npz path at <clip>.single.npz (and use it for the camera reads too — it has L_w2c/ K_leader). Drop the .refs.npz requirement in the eligibility loop and don't open refs.
  • person_pose=False — .single.npz has L_pose (self-only), NOT L_pose_person; with person_pose=True the loader would KeyError.
  • num_reference_frames=0 (R=0) so the reference block is skipped.
  • Everything else is unchanged: _stack_jpeg decode, np.unpackbits(vis_packed) with vis_shape=(77,504,504), _fit_poses(w2c,(77)), _resize to your train resolution — all already match the .single.npz layout.

Verified against a real file (62531cc1_006159.single.npz, 28.5MB): warped/pose/target are JPEG object arrays; vis_shape=(77,504,504) with packed-bit capacity exactly 77*504*504; L_w2c[0]==H_w2c[0]==identity (per-view canonical). So with the three config flips above, the existing loader/decoder path lines up 1:1.


3. Things to know / gotchas

  • Env / launch: sh/train_nymeria_longer.sh, run as CUDA_VISIBLE_DEVICES=0,1,2,3 EXP=<your_exp_name> NPROC=4 bash sh/train_nymeria_longer.sh. Conda env ego_dh. Needs COSMOS_QWEN_TOKENIZER_DIR=/data/cosmos_reason1_7b for the online text encoder, IMAGINAIRE_OUTPUT_ROOT=<out>.
  • Online text encoding: CoMind emits ai_caption; the model text-encodes it online each step (Qwen). Keep the tokenizer dir env set or conditioning will crash on t5_text_embeddings=None.
  • DCP checkpoints are reshardable across NPROC (4↔3↔2 resume works). save_iter in the experiment.
  • register the experiment: append your function to the experiments = [...] list at the bottom of nymeria_pose_2actor.py, and confirm it builds: python -c "..." overriding experiment=<name> (see how the file builds configs) before launching 4-GPU.
  • state_t: latent temporal length = 1 + (T-1)//4. For 77 frames it's 20. It appears in the model config; keep it consistent between data (num_video_frames) and model.
  • AR / noise recipe: predict2/models/video2world_model_rectified_flow.py has an optional noisy_conditioning_* + conditional_frames_probs recipe (teacher-forcing for autoregressive rollout). It is OFF by default ({1:1.0}, prob 0). Ignore it unless you want AR — not needed for the single-ego baseline.
  • Cross-attention: you asked whether to separate it — you don't need to. The DiT already runs unified self-attention over B (V·t) H W D; with no cross-view conditioning the two views just don't share useful signal, which is exactly the single-ego behavior. Leaving it unified is simplest and correct.

4. File map (what's in this bundle)

cosmos_predict2/_src/predict2_multiview/
  configs/vid2vid/experiment/nymeria_pose_2actor.py   # all experiments (copy comind_* -> your single-ego exp)
  configs/vid2vid/defaults/{conditioner,data,model}.py
  datasets/nymeria_pairs.py            # NymeriaActorActorDataset + loaders + cs.store data configs
  datasets/comind_pairs.py             # CoMind dataset/loader
  networks/multiview_pose_dit.py       # the DiT: all the enable_* / view-emb / reference embedders
  models/multiview_pose_model_rectified_flow.py   # conditioning preprocessing (refs/plucker/refpose/depth)
  models/multiview_vid2vid_model_rectified_flow.py
  callbacks/nymeria_validation_viz.py  # validation grid
predict2/models/video2world_model_rectified_flow.py  # base rectified-flow denoise (+optional AR noise)
sh/train_nymeria_longer.sh            # launch

Summary of the diff you need: turn OFF enable_reference_frames, enable_reference_plucker, enable_reference_pose, num_reference_frames=0, shared_reference=False, concat_view_embedding=False; KEEP enable_plucker=True but switch its canonical to each view's own first frame (§1f); from base 2B; and feed CoMind clips rebuilt with own-past warp + own-only pose. Keep warped_cond + pose + Plücker channels and unified cross-view attention.