Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta 26A5353q, 2026-06-11).

Mirror of mlboydaisuke/VJEPA2-ViTL-SSv2-CoreAI β€” the canonical repo (CoreAI Model Zoo). Updates land there first.

This model has no row on DeviceMark, the on-device LLM leaderboard.

V-JEPA 2 (ViT-L, SSv2 action recognition) β€” Apple Core AI

V-JEPA 2 (Meta AI) running natively on the Apple Core AI engine β€” the zoo's first world model: a self-supervised video encoder that learns by predicting in representation space (JEPA), here with the Something-Something v2 action head (174 classes of physical interactions β€” put/lift/push/roll/cover/pretend…).

  • One bundle: ViT-L backbone (3D RoPE attention) + attentive pooler + classifier, ~375M params, fp16 ~675 MB.
  • I/O: pixel_values_videos [1,16,3,256,256] (16 frames, RGB 0..1, ImageNet mean/std) β†’ logits [1,174] (labels.json).
  • Verified: engine vs PyTorch reference cosine 0.999996, top-5 identical; a synthetic motion probe (square moving up vs down) flips the predicted direction correctly.
  • Speed: ~150–180 ms per 16-frame clip on an M4 Max (GPU) β€” real-time video understanding.

Use it

New to Core AI? Start with CoreAIKit 0.7.3. Follow its requirements and first-run steps for qwen3-0.6b, then open the same release's ChatDemo. The README records the tested OS/SDK and download size; model and device coverage is stated per example.

⚑ One line β€” this model is the default behind the kit's task op (import CoreAIOps; no session, no model plumbing, downloads on first use):

let actions = try await CoreAI.recognizeAction(videoAt: videoURL)

Every op, one shape β€” Cookbook.

▢️ Run it (source) β€” the ActionCamera runner (live camera action recognition, one app for every video model in the catalog):

git clone --branch 0.7.3 --depth 1 https://github.com/john-rocky/coreai-kit
export DEVELOPER_DIR=/Applications/Xcode-27.0.0-RC.app/Contents/Developer
open -a /Applications/Xcode-27.0.0-RC.app coreai-kit/Examples/ActionCamera/ActionCamera.xcodeproj
# β†’ Run, then pick "V-JEPA 2 ViT-L (SSv2)" in the model picker

# agents / headless (macOS):
cd coreai-kit/Examples/ActionCamera
swift run -c release action-cli --model vjepa2-vitl-ssv2 --video sample.mp4

Use Xcode build 27A266a from the release's .xcode-pin; adjust the app path if your installation is named differently.

πŸ’» Build with it β€” complete; the glue is kit API, copy-paste runs:

import CoreAIKitVision

let recognizer = try await ActionRecognizer(catalog: "vjepa2-vitl-ssv2")
let actions = try await recognizer.classify(videoAt: videoURL, topK: 3)
// actions: ranked [Prediction] β€” .label ("Pushing [something] from left to right"),
// .probability; 174 SSv2 classes, fully on-device

The take-home is Examples/ActionCamera/Sources/QuickStart.swift β€” this exact code as one typed function, no UI; the CLI is an argument shell over it, and the GUI classifies a rolling 16-frame clip from CameraFeed. Live camera? Keep the last 16 CameraFeed frames and call classify(frames:) β€” other frame counts are uniformly resampled to 16. The bundled sample.mp4 is a synthetic clip (a hand pushing a block); point --video at real footage for real results.

Integration checklist

  • SPM: https://github.com/john-rocky/coreai-kit (exact 0.7.3) β†’ product CoreAIKitVision
  • Info.plist: NSCameraUsageDescription β€” only for the live camera; the snippet needs none
  • Entitlements: none needed
  • First run downloads the model β€” ~708 MB (Mac) / ~710 MB (iPhone) β€” then it loads from the local cache (Application Support; progress via the downloadProgress callback)
  • Measure in Release β€” Debug is ~3Γ— slower on per-token host work

Files

path what
macos/vjepa2_ssv2_fp16.aimodel fp16 bundle (macOS / JIT)
ios/vjepa2_ssv2_fp16.aimodel the same JIT bundle, byte for byte; every iPhone generation specializes it on its first load
ios-h18p/vjepa2_ssv2_fp16.h18p.aimodelc compiled ahead of time for the iPhone 17 Pro (h18p), that phone only; moved from ios/ in revision 672f4a0c (2026-09-26)
macos/labels.json, ios/labels.json, ios-h18p/labels.json 174 SSv2 class names
macos/metadata.json, ios/metadata.json I/O + preprocessing spec

The JIT bundle in ios/ has not been run on an iPhone yet. The iPhone 18 Pro specialized other JIT graphs of up to 1.6 GB on its own. It refuses an h18p bundle with incompatibleCompiledAssetArchitecture (knowledge/jit-distribution.md).

Live demo app: coreai-video β€” camera β†’ live top-3 actions. iPhone 17 Pro with the h18p AOT bundle: ~0.34 s per 16-frame clip.

Preprocessing

Sample 16 frames uniformly from the clip, resize+center-crop to 256Γ—256, scale to 0..1, normalize with ImageNet mean [0.485,0.456,0.406] / std [0.229,0.224,0.225], layout [1,16,3,256,256].

Credits


More models in this format: Core AI Model Zoo β€” 75 models, each with the recipe that produced it.

Want a different model on-device? Open a request β€” free, open weights only; the export and its measured numbers get published publicly.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for coreai-community/VJEPA2-ViTL-SSv2-CoreAI

Quantized
(3)
this model