--- license: apache-2.0 base_model: Mapika/decider-2b-vision base_model_relation: quantized pipeline_tag: image-text-to-text library_name: coreai language: - en tags: - coreai - core-ai - apple - on-device - decision-model - vision-language - multiple-choice --- # decider-2b-vision — Core AI Apple Core AI (`.aimodel`) conversion of [Mapika/decider-2b-vision](https://huggingface.co/Mapika/decider-2b-vision/tree/863e290863655f1d6b69324d77d09ac972d21609) (revision `863e290`, Apache-2.0): an image and lettered questions in, the probability of every option out, from one forward pass. | Variant | Path | Size | Requires | Tested on | |---|---|---:|---|---| | Decoder, int8 with layers 0, 2, 5 fp16 (ship) | `gpu-pipelined/decider_2b_vision_decode_int8mix_pf16/` | 2,664 MB | macOS 27; iOS 27 with the increased-memory-limit entitlement | M4 Max, macOS 27.0 (26A428); iPhone 18 Pro, iOS 27.0 (24A437), 2026-09-29 | | Decoder, fp16 (reference) | `gpu-pipelined/decider_2b_vision_decode_fp16_pf16/` | 3,786 MB | macOS 27 | M4 Max, macOS 27.0 (26A428), 2026-09-29 (not run on the iPhone) | | Vision tower, 256×256 → 64 image rows | `gpu-pipelined/decider_2b_vision_g256_vision_fp16w32/` | 660 MB | macOS 27, iOS 27 | M4 Max; iPhone 18 Pro, 2026-09-29 | | Vision tower, 448×448 → 196 image rows | `gpu-pipelined/decider_2b_vision_g448_vision_fp16w32/` | 663 MB | macOS 27, iOS 27 | M4 Max; iPhone 18 Pro, 2026-09-29 | A decision needs one decoder and one tower. On the iPhone 18 Pro the same files run by device JIT: a `g256` decision in 0.75–0.79 s, `g448` in 1.4 s ([measured below](#iphone-18-pro-ios-270-24a437-device-jit-2026-09-29)). ```swift import CoreAI import DeciderVision let root = URL(filePath: "decider-2b-vision-CoreAI/gpu-pipelined") // a download of the HF repo var decoderOptions = SpecializationOptions(preferredComputeUnitKind: .gpu) decoderOptions.expectFrequentReshapes = true let decider = try await VisionDecider( bundle: root.appending(path: "decider_2b_vision_decode_int8mix_pf16"), towers: [.g256: root.appending(path: "decider_2b_vision_g256_vision_fp16w32/decider_2b_vision_g256_vision_fp16w32.aimodel")], decoderOptions: decoderOptions, towerOptions: SpecializationOptions(preferredComputeUnitKind: .gpu)) let probs = try await decider.decide(image: cgImage /* a CGImage */, context: "This is a visual question about the image.", questions: [DecisionQuestion(text: "How many circles are in the image?", options: ["1", "2", "3", "4", "5"])], grid: .g256) // [[Double]], one row per question ``` The snippet uses the [`DeciderVision`](https://github.com/john-rocky/coreai-model-zoo/tree/main/apps/DeciderVision) Swift package from the model zoo. The zoo card and the gate scripts live in [coreai-model-zoo](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/decider-2b-vision/README.md); the rest of this page is the same text with repository links. Source [Mapika/decider-2b-vision](https://huggingface.co/Mapika/decider-2b-vision/tree/863e290863655f1d6b69324d77d09ac972d21609) (revision `863e290`) · base Qwen/Qwen3.5-2B-Base A **decision model that reads an image**. Give it an image (a photo, a diagram, a game frame), a short context and one or more questions with up to 10 lettered options each. It returns a probability for every option, read from the letter logits at each question's answer slot in one forward pass. It never generates text. Text-only questions go through the same weights. Mapika transplanted the v5 text weights of decider-2b into the Qwen3.5-2B vision-language model and fine-tuned it for one epoch on game frames labelled by scripted policies, multiple-choice image tasks from The Cauldron and a replay of the text mixture, then ran PPO from pixels (the author's card and `decider/vision/` in [Mapika/decider](https://github.com/Mapika/decider)). The text mixture's teacher-written items come from a locally run Qwen3.5-27B (`decider/data/teacher_*.py`). The author's numbers, quoted from the source card and not re-measured here: Visual7W (held out) accuracy 0.89 and ECE 0.03 on 300 items; 0.80–0.95 on six of the mixture's The Cauldron tasks. This port runs that readout as two Core AI graphs. The **vision tower** is baked at a fixed grid: `g256` (a 256×256 tile, 64 image tokens) for game frames and speed, `g448` (448×448, 196 tokens) for photos. It stores fp16 weights and computes in fp32 (660 / 663 MB). The **decoder** is the Qwen3.5 hybrid (18 Gated DeltaNet + 6 full-attention layers) with token ids in and the tower's rows as a static input; it derives the three-plane M-RoPE positions inside the graph. Its linears are int8 per block of 32 except in layers 0, 2 and 5, and the tied embedding / head stays fp16. It is one bundle with an S=1 `main` and an S=16 `prefill` function, 2.64 GB. The gate is probability parity with the author's own fp32 code, on a self-made fixture and on 500 held-out photo runs. ## Readout contract One row per image; every question of the row is answered in the same pass: ``` <|vision_start|><|vision_end|>Context:\n\n\nQuestion 1: \nOptions:\n(A)