Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta 26A5353q, 2026-06-11).
Mirror of
mlboydaisuke/VJEPA2-ViTL-SSv2-CoreAIβ the canonical repo (CoreAI Model Zoo). Updates land there first.
This model has no row on DeviceMark, the on-device LLM leaderboard.
V-JEPA 2 (ViT-L, SSv2 action recognition) β Apple Core AI
V-JEPA 2 (Meta AI) running natively on the Apple Core AI engine β the zoo's first world model: a self-supervised video encoder that learns by predicting in representation space (JEPA), here with the Something-Something v2 action head (174 classes of physical interactions β put/lift/push/roll/cover/pretendβ¦).
- One bundle: ViT-L backbone (3D RoPE attention) + attentive pooler + classifier, ~375M params, fp16 ~675 MB.
- I/O:
pixel_values_videos [1,16,3,256,256](16 frames, RGB 0..1, ImageNet mean/std) βlogits [1,174](labels.json). - Verified: engine vs PyTorch reference cosine 0.999996, top-5 identical; a synthetic motion probe (square moving up vs down) flips the predicted direction correctly.
- Speed: ~150β180 ms per 16-frame clip on an M4 Max (GPU) β real-time video understanding.
Use it
New to Core AI? Start with CoreAIKit 0.7.3. Follow its requirements and first-run steps for qwen3-0.6b, then open the same release's ChatDemo. The README records the tested OS/SDK and download size; model and device coverage is stated per example.
β‘ One line β this model is the default behind the kit's task op
(import CoreAIOps; no session, no model plumbing, downloads on first use):
let actions = try await CoreAI.recognizeAction(videoAt: videoURL)
Every op, one shape β Cookbook.
βΆοΈ Run it (source) β the ActionCamera runner (live camera action recognition, one app for every video model in the catalog):
git clone --branch 0.7.3 --depth 1 https://github.com/john-rocky/coreai-kit
export DEVELOPER_DIR=/Applications/Xcode-27.0.0-RC.app/Contents/Developer
open -a /Applications/Xcode-27.0.0-RC.app coreai-kit/Examples/ActionCamera/ActionCamera.xcodeproj
# β Run, then pick "V-JEPA 2 ViT-L (SSv2)" in the model picker
# agents / headless (macOS):
cd coreai-kit/Examples/ActionCamera
swift run -c release action-cli --model vjepa2-vitl-ssv2 --video sample.mp4
Use Xcode build 27A266a from the release's .xcode-pin; adjust the app path if your installation is named differently.
π» Build with it β complete; the glue is kit API, copy-paste runs:
import CoreAIKitVision
let recognizer = try await ActionRecognizer(catalog: "vjepa2-vitl-ssv2")
let actions = try await recognizer.classify(videoAt: videoURL, topK: 3)
// actions: ranked [Prediction] β .label ("Pushing [something] from left to right"),
// .probability; 174 SSv2 classes, fully on-device
The take-home is Examples/ActionCamera/Sources/QuickStart.swift
β this exact code as one typed function, no UI; the CLI is an argument shell over it, and
the GUI classifies a rolling 16-frame clip from CameraFeed.
Live camera? Keep the last 16 CameraFeed frames and call classify(frames:) β other
frame counts are uniformly resampled to 16. The bundled sample.mp4 is a synthetic
clip (a hand pushing a block); point --video at real footage for real results.
Integration checklist
- SPM:
https://github.com/john-rocky/coreai-kit(exact 0.7.3) β product CoreAIKitVision - Info.plist:
NSCameraUsageDescriptionβ only for the live camera; the snippet needs none - Entitlements: none needed
- First run downloads the model β ~708 MB (Mac) / ~710 MB (iPhone) β then it loads from the
local cache (Application Support; progress via the
downloadProgresscallback) - Measure in Release β Debug is ~3Γ slower on per-token host work
Files
| path | what |
|---|---|
macos/vjepa2_ssv2_fp16.aimodel |
fp16 bundle (macOS / JIT) |
ios/vjepa2_ssv2_fp16.aimodel |
the same JIT bundle, byte for byte; every iPhone generation specializes it on its first load |
ios-h18p/vjepa2_ssv2_fp16.h18p.aimodelc |
compiled ahead of time for the iPhone 17 Pro (h18p), that phone only; moved from ios/ in revision 672f4a0c (2026-09-26) |
macos/labels.json, ios/labels.json, ios-h18p/labels.json |
174 SSv2 class names |
macos/metadata.json, ios/metadata.json |
I/O + preprocessing spec |
The JIT bundle in ios/ has not been run on an iPhone yet. The iPhone 18 Pro specialized other JIT
graphs of up to 1.6 GB on its own. It refuses an h18p bundle with
incompatibleCompiledAssetArchitecture
(knowledge/jit-distribution.md).
Live demo app: coreai-video β camera β live top-3 actions. iPhone 17 Pro with the h18p AOT bundle: ~0.34 s per 16-frame clip.
Preprocessing
Sample 16 frames uniformly from the clip, resize+center-crop to 256Γ256, scale to 0..1, normalize
with ImageNet mean [0.485,0.456,0.406] / std [0.229,0.224,0.225], layout [1,16,3,256,256].
Credits
- Meta AI β V-JEPA 2 (MIT).
- Conversion + Core AI port: coreai-model-zoo.
More models in this format: Core AI Model Zoo β 75 models, each with the recipe that produced it.
Want a different model on-device? Open a request β free, open weights only; the export and its measured numbers get published publicly.
- Downloads last month
- -
Model tree for coreai-community/VJEPA2-ViTL-SSv2-CoreAI
Base model
facebook/vjepa2-vitl-fpc64-256