twanghcmut's picture
|
download
raw
5.25 kB

MoGe-2 monocular depth — qualitative previews

Full-length monocular depth over the clips published in this bucket, rendered as a three-panel video: RGB | depth | normal, each panel at the clip's native 1280x720.

These are previews for judging detail by eye. They are not a dataset tier and nothing in the pipeline consumes them — the pipeline's own depth comes from PointWorld scene_flows (sparse metric 3D tracks), not from MoGe.

Model

Ruicheng/moge-2-vitl-normal, resolution_level=9.

resolution_level=9 is the default and already maps to the top of MoGe-2's num_tokens_range (3600 tokens) — the finest setting the model offers.

MoGe-3 is not used, because it does not exist as runnable code. The paper (arXiv 2607.17967, "Fine-Detail Monocular Geometry Estimation…") was announced in the upstream README on 2026-07-21, but at pinned commit 925b8ed — still origin/main as of this run — there is no moge/model/v3.py and the pretrained table lists moge-3-vitg/moge-3-vitl as "coming soon" with no HF repo id.

How to read the panels

  • Depth uses one global normalisation for the whole clip (robust 2/98 percentiles over every valid pixel of every frame), never per frame. Per-frame normalisation makes the panel breathe as the near/far extremes move, which by eye is indistinguishable from the model's own temporal instability — the exact thing these videos exist to let someone judge. The colour range in metres is burned into the panel label.
  • Normal is where fine detail actually shows. Depth is dominated by the metre-scale front-to-back ramp of the workbench; the normal map spends its whole colour budget on surface relief instead.
  • No fov_x was passed, so MoGe estimates its own FOV per frame even though this project knows the true DROID intrinsics. That is the out-of-the-box path — priming the model with the calibration would measure a best case that arbitrary video would not get.
  • Videos are muxed at 15 fps, the DROID trajectory rate. The source mp4 containers advertise 60/1, but their frames are the 15 Hz trajectory samples (this repo indexes poses per video frame at DEFAULT_TRAJECTORY_FPS = 15.0). Playing at 60 would look like temporal jitter that is not there.

Clips

file camera frames depth p2–p98 (m) size
AUTOLab+0d4edc83+2023-10-21-19h-06m-10s__22008760.mp4 ext1 178 0.31–1.41 20.7 MB
AUTOLab+0d4edc83+2023-10-21-19h-07m-04s__22008760.mp4 ext1 127 0.27–1.48 15.3 MB
AUTOLab+0d4edc83+2023-10-21-19h-07m-04s__24400334.mp4 ext2 127 0.42–1.65 15.5 MB
AUTOLab+0d4edc83+2023-10-21-19h-07m-18s__22008760.mp4 ext1 146 0.25–1.46 16.9 MB
AUTOLab+0d4edc83+2023-10-21-19h-08m-08s__22008760.mp4 ext1 148 0.26–1.48 17.7 MB
AUTOLab+0d4edc83+2023-10-21-19h-08m-25s__22008760.mp4 ext1 133 0.27–1.45 16.0 MB
AUTOLab+0d4edc83+2023-10-21-19h-08m-41s__22008760.mp4 ext1 165 0.28–1.46 18.8 MB
AUTOLab+0d4edc83+2023-10-21-19h-08m-58s__22008760.mp4 ext1 198 0.19–1.36 23.0 MB
AUTOLab+0d4edc83+2023-10-21-19h-09m-18s__22008760.mp4 ext1 124 0.27–1.42 15.0 MB
AUTOLab+0d4edc83+2023-10-21-19h-09m-53s__22008760.mp4 ext1 175 0.29–1.49 19.3 MB

All of these are the same scene (AUTOLab+0d4edc83, "Put brick in drawer shelf and close drawer"), recorded seconds apart, because that is the whole of what this bucket holds. Ten clips agreeing therefore is not evidence that MoGe-2 is stable — it is one condition measured ten times. The only genuinely different viewpoint in the set is camera 24400334 (ext2).

object_step.json — the quantitative companion

Whether the model separates the manipulated object from the surface under it, measured with the pipeline's own object mask (master/prompt_0_masks.h5):

  • step_mm — mean depth of a 25 px ring around the object minus mean depth inside it. Positive = object in front of its surroundings.
  • edge_ratio — median |∇depth| on the mask boundary over the median in the ring. 1.0 = boundary indistinguishable from flat surface. This is the half of the question a step size cannot answer: a model that smears a real 20 mm step over 30 px still reports 20 mm.

Both are computed within a single frame on purpose. With no fov_x given, MoGe estimates its own scale per frame and that scale drifts, so absolute metres are not comparable across frames; an inside-minus-outside difference cancels a per-frame scale factor to first order and is.

Two caveats that travel with these numbers:

  • step_mm_p10 goes negative on several clips. That is the measurement's limit, not the model's error: once the gripper closes on the object, the 25 px ring contains the fingers, which are nearer the camera than the object, so the difference flips sign. The median is robust to this; the p10 is not.
  • Two clips have no prompt_0_masks.h5 at all (...07m-04s/24400334, ...09m-53s/22008760) — both are the entries with n_win=0 in the dataset manifest, i.e. datagen never produced object masks for them. They ship a video and no step.

Xet Storage Details

Size:
5.25 kB
·
Xet hash:
53057f67889b173e4c4d6725555cb66e1d4fcf9f307aabb2ffd73542c3012076

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.