Download README.md from dhyun22/train_code_for_hs: direct link, hf CLI and curl.
- Browser
- Download file 12.1 kB
-
https://huggingface.co/dhyun22/train_code_for_hs/resolve/main/README.md
- Command line
-
hf download hf://dhyun22/train_code_for_hs/README.md
-
curl -L -o README.md https://huggingface.co/dhyun22/train_code_for_hs/resolve/main/README.md
Training code — CoMind 2-view generation (single-ego setup)
This bundle is our current 2-view egocentric video training stack (Cosmos-Predict2.5, predict2_multiview).
It is set up for a 2-view actor-actor model with a lot of cross-view conditioning. You are going to
strip it down to a single-ego generation setup that still outputs 2 views, trained from scratch on the
CoMind dataset. This README lists exactly what to change and what to watch out for.
0. What the current model conditions on (the parts you will REMOVE / keep)
Per view, each generated clip is conditioned on:
| Signal | Key | What it is | Single-ego plan |
|---|---|---|---|
| warped RGB | control_input_warped |
RGB warped into the target view from a combined pool (both views' past) | KEEP the channel, but change the DATA: warp only from that view's own past context |
| pose | control_input_pose |
skeletons. person_pose=True = identity-colored, both people visible |
KEEP the channel, change the DATA: render only the wearer's own pose |
| in-context references | reference_frames (+ reference_pose, reference_cam_w2c) |
clean appearance/pose anchor frames appended as extra tokens | REMOVE entirely |
| camera Plücker | plucker_map |
per-pixel rays. Currently in a shared cross-view canonical frame (view0 frame0) | KEEP (enable_plucker=True), but change the canonical to each view's OWN first frame — see §1f |
| reference Plücker | reference_plucker_map |
posed refs' rays | REMOVE (enable_reference_plucker=False; no refs) |
| view embedding | net view_embeddings |
per-view identity added to tokens | REMOVE (drop concat_view_embedding) |
| cross-view self-attention | token layout B (V·t) … |
both views attend each other | Leave as-is — you do NOT need to separate it; unified attention is fine, and with all cross-view conditioning removed the two views are effectively independent anyway |
So the single-ego model keeps: warped_cond (own past) + own pose + camera Plücker (own-first-frame canonical), generated per view, no refs / no reference-Plücker / no view-embedding, trained from the Cosmos base 2B.
1. Code changes (all in cosmos_predict2/_src/predict2_multiview/)
Make a new experiment function in
configs/vid2vid/experiment/nymeria_pose_2actor.py (copy comind_actoractor_personpose_shared as a start),
and flip these toggles. The flags exist already — you're just turning conditioning OFF.
1a. model config (models/multiview_pose_model_rectified_flow.py fields, set in the experiment)
cfg["model"]["config"]["num_reference_frames"] = 0 # was 4
cfg["model"]["config"]["enable_reference_plucker"] = False # was True (no refs)
cfg["model"]["config"]["enable_reference_pose"] = False # was True (no refs)
cfg["model"]["config"]["enable_plucker"] = True # KEEP camera Plücker (own-first-frame canonical, §1f)
1b. net config (networks/multiview_pose_dit.py, set via cfg["model"]["config"]["net"].update(...))
cfg["model"]["config"]["net"].update(
enable_reference_frames = False,
num_reference_frames = 0,
shared_reference = False,
enable_reference_plucker= False,
enable_reference_pose = False,
enable_plucker = True, # KEEP Plücker embedder (6-ch ray map -> zero-init add)
concat_view_embedding = False, # <-- drops the view embedding
)
pose_mode stays "vae_concat" (pose is VAE-encoded and concatenated into cond_embedder; keep it). The
cond_embedder in-channels auto-compute as warped_latent(16) + visibility(1) + pose_latent(16) — leave that.
1f. Plücker canonical frame = each view's OWN first frame — already baked in the data, NO code change
We want each view's Plücker rays expressed relative to its OWN first frame (not a shared cross-view frame).
The .single.npz data already bakes this: L_w2c[0] and H_w2c[0] are the identity matrix
(cam_frame: per_view_canonical_frame0), i.e. every view's camera poses are already canonicalized to its own
frame0. The model's existing _preprocess (~line 185) computes c2w_ref = inv(w2c[:, 0, 0]); since view0
frame0 is identity, c2w_ref = I and the shared transform is a no-op, so each view's rays come out in its
own already-canonical frame — exactly what we want. Therefore: keep the model Plücker code AS-IS. Do NOT add
a per-view-canonical loop — the data already did it (adding one would be redundant). Just keep
enable_plucker=True, enable_reference_plucker=False, and drop the reference_plucker_map branch (refs off).
1c. checkpoint = from scratch
cfg["checkpoint"]["load_path"] = _base_2b_multiview_ckpt() # Cosmos base 2B, NOT a warm-start
cfg["checkpoint"]["strict_resume"] = False
Because refs/Plücker/view-emb are gone and the pose/warp embedders are zero-init or freshly-shaped, iter-1 loss will start HIGH (~base video prior) — that's correct for a from-scratch run, unlike a warm-start.
1d. data loader
Point override /data_train and override /data_val at a CoMind data config that:
- reads your regenerated CoMind clips (own-past warp + own-only pose),
- sets
num_reference_frames=0,shared_reference=False,emit_reference_pose=False(do NOT emit anyreference_*),person_pose=your choice (you're rendering own-only pose, so the identity-color flag is moot).
The CoMind loader lives in datasets/comind_pairs.py / the actor-actor path in datasets/nymeria_pairs.py
(get_nymeria_actor_actor_loader, class NymeriaActorActorDataset). Register a new cs.store(...) data
config with the flags above.
1e. validation viz
callbacks/nymeria_validation_viz.py: set num_reference_frames=0, shared_reference=False, emit_reference_pose=False in the viz kwargs too, or the viz dataset build will look for refs that aren't there.
2. DATA — the ablation clips are the sibling .single.npz (already being generated)
The single-ego ablation data is stored as a sibling <clip>.single.npz next to the untouched originals
(<clip>.npz, .campose.npz, .refs.npz). The .single.npz is self-contained for training and its keys
match the loader's existing convention. Contents (JPEG byte-arrays for image streams, same as the originals):
L_/H_target GT video (JPEG) -> video
L_/H_warped OWN-only warp (srcview==0 kept; ==1 -> black hole) -> control_input_warped
L_/H_pose SELF-only skeleton (own hands only) -> control_input_pose
L_/H_vis_packed + vis_shape packed visibility (hole frames vis=False) -> control_input_visibility
L_/H_w2c (77,4,4) PER-VIEW canonical (each view frame0 = identity) -> camera_w2c
K_leader/K_helper (3,3) intrinsics -> camera_K
L_/H_srcf source frame idx (hole=-1) [loader ignores]
warp_mode/pose_mode/cam_frame/n_hole_* markers [loader ignores]
IMPORTANT — source strictly from .single.npz. The original clip.npz ALSO has L_warped/L_pose/ L_vis_packed, but those are the OLD combined-pool warp + both-people pose. Read warped/pose/vis (and w2c/K)
only from .single.npz. Since .single.npz also carries L_target, the simplest wiring is to use
.single.npz as the sole source and not open clip.npz at all (keypoints aren't read by training). No
.refs.npz (references are off).
2a. Loader adaptation (small — the format already matches)
Start from datasets/comind_pairs.py (ComindActorActorDataset), which already reads L_/H_ keys as JPEG
streams. Change:
- Point the npz path at
<clip>.single.npz(and use it for the camera reads too — it hasL_w2c/K_leader). Drop the.refs.npzrequirement in the eligibility loop and don't open refs. person_pose=False—.single.npzhasL_pose(self-only), NOTL_pose_person; withperson_pose=Truethe loader would KeyError.num_reference_frames=0(R=0) so the reference block is skipped.- Everything else is unchanged:
_stack_jpegdecode,np.unpackbits(vis_packed)withvis_shape=(77,504,504),_fit_poses(w2c,(77)),_resizeto your train resolution — all already match the.single.npzlayout.
Verified against a real file (62531cc1_006159.single.npz, 28.5MB): warped/pose/target are JPEG object arrays;
vis_shape=(77,504,504) with packed-bit capacity exactly 77*504*504; L_w2c[0]==H_w2c[0]==identity
(per-view canonical). So with the three config flips above, the existing loader/decoder path lines up 1:1.
3. Things to know / gotchas
- Env / launch:
sh/train_nymeria_longer.sh, run asCUDA_VISIBLE_DEVICES=0,1,2,3 EXP=<your_exp_name> NPROC=4 bash sh/train_nymeria_longer.sh. Conda envego_dh. NeedsCOSMOS_QWEN_TOKENIZER_DIR=/data/cosmos_reason1_7bfor the online text encoder,IMAGINAIRE_OUTPUT_ROOT=<out>. - Online text encoding: CoMind emits
ai_caption; the model text-encodes it online each step (Qwen). Keep the tokenizer dir env set or conditioning will crash ont5_text_embeddings=None. - DCP checkpoints are reshardable across NPROC (4↔3↔2 resume works).
save_iterin the experiment. - register the experiment: append your function to the
experiments = [...]list at the bottom ofnymeria_pose_2actor.py, and confirm it builds:python -c "..."overridingexperiment=<name>(see how the file builds configs) before launching 4-GPU. state_t: latent temporal length =1 + (T-1)//4. For 77 frames it's 20. It appears in the model config; keep it consistent between data (num_video_frames) and model.- AR / noise recipe:
predict2/models/video2world_model_rectified_flow.pyhas an optionalnoisy_conditioning_*+conditional_frames_probsrecipe (teacher-forcing for autoregressive rollout). It is OFF by default ({1:1.0}, prob 0). Ignore it unless you want AR — not needed for the single-ego baseline. - Cross-attention: you asked whether to separate it — you don't need to. The DiT already runs unified
self-attention over
B (V·t) H W D; with no cross-view conditioning the two views just don't share useful signal, which is exactly the single-ego behavior. Leaving it unified is simplest and correct.
4. File map (what's in this bundle)
cosmos_predict2/_src/predict2_multiview/
configs/vid2vid/experiment/nymeria_pose_2actor.py # all experiments (copy comind_* -> your single-ego exp)
configs/vid2vid/defaults/{conditioner,data,model}.py
datasets/nymeria_pairs.py # NymeriaActorActorDataset + loaders + cs.store data configs
datasets/comind_pairs.py # CoMind dataset/loader
networks/multiview_pose_dit.py # the DiT: all the enable_* / view-emb / reference embedders
models/multiview_pose_model_rectified_flow.py # conditioning preprocessing (refs/plucker/refpose/depth)
models/multiview_vid2vid_model_rectified_flow.py
callbacks/nymeria_validation_viz.py # validation grid
predict2/models/video2world_model_rectified_flow.py # base rectified-flow denoise (+optional AR noise)
sh/train_nymeria_longer.sh # launch
Summary of the diff you need: turn OFF enable_reference_frames, enable_reference_plucker,
enable_reference_pose, num_reference_frames=0, shared_reference=False, concat_view_embedding=False;
KEEP enable_plucker=True but switch its canonical to each view's own first frame (§1f); from base 2B;
and feed CoMind clips rebuilt with own-past warp + own-only pose. Keep warped_cond + pose + Plücker
channels and unified cross-view attention.