Depth Anything 3 (small+base) Core AI bundles + card

Browse files

Files changed (14) hide show

.gitattributes +4 -0
README.md +84 -0
base/da3-base_float16.aimodel/main.hash +1 -0
base/da3-base_float16.aimodel/main.mlirb +3 -0
base/da3-base_float16.aimodel/metadata.json +7 -0
base/da3-base_float32.aimodel/main.hash +1 -0
base/da3-base_float32.aimodel/main.mlirb +3 -0
base/da3-base_float32.aimodel/metadata.json +7 -0
small/da3-small_float16.aimodel/main.hash +2 -0
small/da3-small_float16.aimodel/main.mlirb +3 -0
small/da3-small_float16.aimodel/metadata.json +7 -0
small/da3-small_float32.aimodel/main.hash +1 -0
small/da3-small_float32.aimodel/main.mlirb +3 -0
small/da3-small_float32.aimodel/metadata.json +7 -0

.gitattributes CHANGED Viewed

@@ -33,3 +33,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+base/da3-base_float16.aimodel/main.mlirb filter=lfs diff=lfs merge=lfs -text
+base/da3-base_float32.aimodel/main.mlirb filter=lfs diff=lfs merge=lfs -text
+small/da3-small_float16.aimodel/main.mlirb filter=lfs diff=lfs merge=lfs -text
+small/da3-small_float32.aimodel/main.mlirb filter=lfs diff=lfs merge=lfs -text

README.md ADDED Viewed

	@@ -0,0 +1,84 @@

+---
+license: apache-2.0
+tags:
+- depth-estimation
+- monocular-depth
+- core-ai
+- coreai
+- apple
+- on-device
+- depth-anything
+pipeline_tag: depth-estimation
+base_model:
+- depth-anything/DA3-SMALL
+- depth-anything/DA3-BASE
+library_name: coreai
+---
+# Depth Anything 3 — Core AI
+**The [coreai-model-zoo](https://github.com/john-rocky/coreai-model-zoo)'s first depth model.**
+Monocular (single-image) **relative depth** estimation running fully on-device on Apple's Core AI
+runtime, as a single static `.aimodel`. A conversion of ByteDance's
+[Depth Anything 3](https://github.com/ByteDance-Seed/depth-anything-3)
+([`depth-anything/DA3-SMALL`](https://huggingface.co/depth-anything/DA3-SMALL) /
+[`DA3-BASE`](https://huggingface.co/depth-anything/DA3-BASE), Apache-2.0): a DINOv2 ViT backbone +
+DPT-style head. Drop in an RGB image, get a depth map (and a confidence map). No NMS, no sampling —
+host post-processing is just a colormap.
+## Bundles
+| dir | variant | params | dtype | size | M4 Max GPU |
+|---|---|---|---|---|---|
+| `small/da3-small_float16.aimodel` | ViT-S | 34.3M | fp16 | **54 MB** | **65.7 FPS** |
+| `small/da3-small_float32.aimodel` | ViT-S | 34.3M | fp32 | 105 MB | 56.5 FPS |
+| `base/da3-base_float16.aimodel` | ViT-B | 135.4M | fp16 | 202 MB | 26.5 FPS |
+| `base/da3-base_float32.aimodel` | ViT-B | 135.4M | fp32 | 402 MB | 23.0 FPS |
+`small · fp16` is the on-device hero — 54 MB, 65 FPS at 504² on an M4 Max, comfortably real-time on
+iPhone-class GPUs. Each `.aimodel` is a directory bundle (`main.mlirb` + `metadata.json`).
+## I/O contract
+```
+input : image [1, 3, 504, 504]  RGB, raw [0, 1]   (ImageNet normalization is folded into the graph)
+output: depth      [1, 504, 504]  relative depth (exp-activated; larger = nearer)
+        depth_conf [1, 504, 504]  confidence
+```
+Host: resize the RGB image to 504 × 504 (e.g. cv2 `INTER_AREA`), feed raw [0, 1], run, then resize
+the depth map back to the original H × W. For display, the DA3 convention is inverse-depth →
+percentile 2–98 normalize → `Spectral` colormap.
+## Fidelity
+- **Bit-exact conversion:** the Core AI engine matches the PyTorch reference at **cos 1.000000** (≤
+  ~1e-5 / ~1e-2 per-pixel for fp32 / fp16) on both CPU and GPU, at any fixed input shape.
+- **vs the official DA3 viewer:** **mean Pearson r ≈ 0.98** across diverse aspect ratios (square
+  inputs r = 1.000) — within DA3's own resolution sensitivity (its 504-vs-518 outputs differ by
+  r ≈ 0.975–0.984).
+## Usage (CoreAIKit / coreai.runtime)
+```python
+import coreai.runtime as rt, numpy as np
+from PIL import Image
+m = await rt.AIModel.load("small/da3-small_float16.aimodel",
+        rt.SpecializationOptions.from_preferred_compute_unit_kind(rt.ComputeUnitKind.gpu()))
+fn = m.load_function("main")
+img = np.asarray(Image.open("photo.jpg").convert("RGB").resize((504, 504)))
+x = (img.astype(np.float16) / 255.0).transpose(2, 0, 1)[None]   # raw [0,1], NCHW
+depth = (await fn({"image": rt.NDArray(x)}))["depth"].numpy().reshape(504, 504)
+```
+## Links
+- Conversion script + model card: [coreai-model-zoo `zoo/depth-anything-3.md`](https://github.com/john-rocky/coreai-model-zoo/blob/main/zoo/depth-anything-3.md)
+- Source: [Depth Anything 3](https://github.com/ByteDance-Seed/depth-anything-3) · Apache-2.0
+---
+*On-device ML / Core ML / Core AI model porting — get in touch: open an issue on the
+[zoo](https://github.com/john-rocky/coreai-model-zoo).*

base/da3-base_float16.aimodel/main.hash ADDED Viewed

	@@ -0,0 +1 @@


1	+ �8\�I��V!��K�פѪ��J�P��Gg��

base/da3-base_float16.aimodel/main.mlirb ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:ed385ced49b3e2975621b8d54bead7a410d1aa19d818bc4ac150a1c04767f7bb
+size 201980350

base/da3-base_float16.aimodel/metadata.json ADDED Viewed

	@@ -0,0 +1,7 @@

+{
+  "license" : "Apache-2.0",
+  "creationDate" : "20260623T154004Z",
+  "assetVersion" : "2.0",
+  "description" : "Depth Anything 3 (DINOv2 ViT-S backbone + DualDPT head), monocular depth. Input: RGB [0,1]; outputs: relative depth + confidence map. https:\/\/github.com\/ByteDance-Seed\/depth-anything-3",
+  "author" : "ByteDance (Depth Anything 3); Core AI export: coreai-model-zoo"
+}

base/da3-base_float32.aimodel/main.hash ADDED Viewed

	@@ -0,0 +1 @@


1	+ ��Z��,�]-T#�Uo��ZC�Ƀu .p

base/da3-base_float32.aimodel/main.mlirb ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8896cd06d05ae7e12ca85d2d5423d7556ff0f8c15a43c10b07c983751f092e70
+size 401763991

base/da3-base_float32.aimodel/metadata.json ADDED Viewed

	@@ -0,0 +1,7 @@

+{
+  "creationDate" : "20260623T153804Z",
+  "license" : "Apache-2.0",
+  "assetVersion" : "2.0",
+  "description" : "Depth Anything 3 (DINOv2 ViT-S backbone + DualDPT head), monocular depth. Input: RGB [0,1]; outputs: relative depth + confidence map. https:\/\/github.com\/ByteDance-Seed\/depth-anything-3",
+  "author" : "ByteDance (Depth Anything 3); Core AI export: coreai-model-zoo"
+}

small/da3-small_float16.aimodel/main.hash ADDED Viewed

	@@ -0,0 +1,2 @@


1	+ `Iց�
2	+ iLm�&��I�?NP`�q�[��4&�

small/da3-small_float16.aimodel/main.mlirb ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6049d6818a0a694c6ded26b38849ae3f4e5060ef71171f975bb29cfa3426e307
+size 54518253

small/da3-small_float16.aimodel/metadata.json ADDED Viewed

	@@ -0,0 +1,7 @@

+{
+  "author" : "ByteDance (Depth Anything 3); Core AI export: coreai-model-zoo",
+  "license" : "Apache-2.0",
+  "assetVersion" : "2.0",
+  "creationDate" : "20260623T153758Z",
+  "description" : "Depth Anything 3 (DINOv2 ViT-S backbone + DualDPT head), monocular depth. Input: RGB [0,1]; outputs: relative depth + confidence map. https:\/\/github.com\/ByteDance-Seed\/depth-anything-3"
+}

small/da3-small_float32.aimodel/main.hash ADDED Viewed

	@@ -0,0 +1 @@


1	+ ��fFhdG�%�Rb�[͵n�]�W[o��T� ?�

small/da3-small_float32.aimodel/main.mlirb ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:1aaaa46646686447e525b60752628a5bcdb56ecb5dd3575b6ffcb354f3093fab
+size 104852846

small/da3-small_float32.aimodel/metadata.json ADDED Viewed

	@@ -0,0 +1,7 @@

+{
+  "description" : "Depth Anything 3 (DINOv2 ViT-S backbone + DualDPT head), monocular depth. Input: RGB [0,1]; outputs: relative depth + confidence map. https:\/\/github.com\/ByteDance-Seed\/depth-anything-3",
+  "author" : "ByteDance (Depth Anything 3); Core AI export: coreai-model-zoo",
+  "creationDate" : "20260623T152907Z",
+  "license" : "Apache-2.0",
+  "assetVersion" : "2.0"
+}