Card: base_model_relation: quantized (so the conversion lists under the base model's Quantizations, not Finetunes)
daec7d8 verified | license: apache-2.0 | |
| tags: | |
| - executorch | |
| - xnnpack | |
| - pte | |
| - on-device | |
| - object-detection | |
| - zero-shot-object-detection | |
| base_model: | |
| - IDEA-Research/grounding-dino-tiny | |
| base_model_relation: quantized | |
| # Grounding DINO tiny β ExecuTorch (open-vocabulary detection) | |
| You name the thing in text and it finds it. No fixed class list: the prompt is | |
| "a person. a car. a bicycle. a dog." today and "a fire hydrant. a traffic cone." | |
| on the next call, with the same weights. | |
| ``` | |
| (pixel_values (1,3,640,640), input_ids (1,32), token_type_ids (1,32), | |
| attention_mask (1,32), text_masks (1,32,32), position_ids (1,32)) | |
| -> scores (1,900,256), boxes (1,900,4) | |
| ``` | |
| | build | file | MB | scores corr | boxes corr | detections agreeing | Mac ms* | | |
| |---|---|---|---|---|---|---| | |
| | fp32 | `grounding_dino_tiny_xnnpack_fp32.pte` | 690.3 | 0.999833 | 0.998992 | **26/26** | 584.1 | | |
| | int8 (dynamic) | `grounding_dino_tiny_xnnpack_int8.pte` | **254.4** | 0.804530 | 0.777331 | **26/27** | 738.7 | | |
| \*Mac arm64, single process, median of 5 β a reference point for relative cost, not a | |
| device number. Torch eager fp32 on the same machine: 376.7 ms. XNNPACK delegate coverage | |
| 63.0% (3503/5559 ops) for fp32, 71.5% for int8. | |
| **Read that int8 row carefully.** Its correlations are 0.80 and 0.78, which on any other | |
| model on this shelf would withdraw the build. Here they are dominated by the queries that | |
| detect nothing, whose coordinates are free to move without changing an answer β and the | |
| detections themselves hold: all 26 that eager finds are found, at IoU > 0.5 with the same | |
| prompt token, plus one box eager does not have. It is 37% of the fp32 file and it is what | |
| you would put on a phone, but the evidence behind it is 26 detections over five | |
| photographs, not a COCO run, and one of those 27 boxes already disagrees. It is also | |
| slower than fp32 on this Mac (738.7 ms against 584.1). | |
| **There is no Core ML build.** The graph reshapes to rank 6 and Core ML's ceiling is 5 β | |
| the same wall RAFT-small hits with its correlation volume. This one is XNNPACK only, and | |
| the same file runs on Android and on the Mac. | |
| ## Running it | |
| **1. The image.** RGB, divide by 255, ImageNet normalise (mean .485/.456/.406, std | |
| .229/.224/.225), resize to 640Γ640. | |
| **2. The prompt.** Lower-case phrases separated by `". "`, ending in a period. Tokenise | |
| with the repo's tokenizer at `padding="max_length", max_length=32`. `input_ids`, | |
| `token_type_ids` and `attention_mask` are int64 `(1,32)`. | |
| **3. The two text tensors that are not the tokenizer's.** `text_masks` and `position_ids` | |
| come from transformers' `generate_masks_with_special_tokens_and_transfer_map(input_ids)` | |
| β the mask that keeps each phrase's tokens attending only to their own phrase. It is | |
| built with `cummax`/`cummin`, which are not in the Core ATen opset, so it is lifted out of | |
| the graph and passed in β `text_masks` is **bool** `(1,32,32)` and `position_ids` is int64 | |
| `(1,32)`, and ExecuTorch will not coerce either. Both depend on `input_ids` alone, never | |
| on the image, so compute them once when the user types the prompt. The model stays | |
| open-vocabulary; only the token count is fixed. | |
| **4. Reading the answer.** `boxes` is `(cx, cy, w, h)` normalised to the input square. | |
| `scores[0, q, t]` is how much query `q` points at text position `t`. A detection is a | |
| query whose best score clears your threshold (0.3 is a reasonable default), and its label | |
| is the `input_ids` token at that argmax. | |
| Two things about that text axis. It is always 256 wide β the model's `max_text_len` β no | |
| matter how long your prompt is, and every column past your prompt's tokens is exactly 0. | |
| And the peak often lands on an article rather than the noun: a query for a person may | |
| argmax on the `a` of "a person". Map the winning token back to the **phrase** it belongs | |
| to, the way transformers' own post-processor does, rather than reading the single token | |
| as the label. | |
| ## What is measured | |
| The number that decides is detections, not correlation. Against eager PyTorch on five | |
| street photographs at threshold 0.3: | |
| **fp32 reproduces 26 of 26**, each at IoU > 0.5 with the same prompt token, and invents | |
| nothing. That includes a photo with 17 detections in it. **int8 scores 26 of 27** β it | |
| finds all of eager's and one box besides. The denominator is whichever build found more, | |
| so an invented detection costs the same as a missed one; counting only the reference's | |
| detections would give a build that returns every box in the image a perfect score. | |
| ```bash | |
| python convert/audit_int8.py grounding_dino_tiny --variant fp32 # or int8 | |
| python convert/verify_cards.py grounding_dino_tiny | |
| ``` | |
| Correlation is in the table above for completeness, and the boxes figure (0.998992) is | |
| below this shelf's usual fp32 bar of 0.999. It is not measuring what it looks like: the | |
| largest box differences are on queries that fire at nothing, where a coordinate is free to | |
| wander without changing any answer. The detection count is what was gated on. | |
| ### The number that had this model filed as broken | |
| Grounding DINO writes **-inf** into every text column the prompt does not use. The score | |
| tensor is 900 queries Γ 256 columns; a four-phrase prompt uses 14 of them, so **94.5% of | |
| the raw logits are -inf**. Correlation over that tensor is undefined, and subtracting two | |
| builds gives NaN rather than a difference. Measured that way this conversion read "corr | |
| 0.857, returns the stopword 'a' where eager finds four people" and sat parked. Both halves | |
| were artefacts: eager returns that same `a` query at the same 0.385, and the detections | |
| were identical all along. | |
| The shipped graph ends in a `sigmoid`, so those positions are exactly 0, the output is in | |
| the 0..1 an app wants, and the gate measures the model instead of its padding. | |
| - **Source**: [IDEA-Research/grounding-dino-tiny](https://huggingface.co/IDEA-Research/grounding-dino-tiny) | |
| - **License**: Apache-2.0 | |
| torch.export -> to_edge_transform_and_lower(partitioner) -> .pte | |
| (conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models)) | |