Card: base_model_relation: quantized (so the conversion lists under the base model's Quantizations, not Finetunes)
dc2af82 verified |
Download README.md from mlboydaisuke/multilingual-e5-base-ExecuTorch: direct link, hf CLI and curl.
- Browser
- Download file 4.71 kB
-
https://huggingface.co/mlboydaisuke/multilingual-e5-base-ExecuTorch/resolve/main/README.md
- Command line
-
hf download hf://mlboydaisuke/multilingual-e5-base-ExecuTorch/README.md
-
curl -L -o README.md https://huggingface.co/mlboydaisuke/multilingual-e5-base-ExecuTorch/resolve/main/README.md
4.71 kB
| license: mit | |
| tags: | |
| - executorch | |
| - xnnpack | |
| - pte | |
| - on-device | |
| - sentence-similarity | |
| - feature-extraction | |
| base_model: | |
| - intfloat/multilingual-e5-base | |
| base_model_relation: quantized | |
| # multilingual-e5-base β ExecuTorch | |
| Multilingual E5, base-sized β one index over 100 languages. Text in, one | |
| 768-dimensional vector out, for search and retrieval that never leaves the | |
| device. | |
| - **Source**: intfloat/multilingual-e5-base β 12 layers, 768 dimensions, 250,002 vocabulary | |
| - **License**: mit | |
| - **Input**: `input_ids` and `attention_mask`, both `[1, 256]` int64 | |
| - **Output**: `[1, 768]`, mean-pooled and L2-normalised inside the graph | |
| ## The recipe is in the graph, and it was read off this repo | |
| sentence-transformers stores it per model, and the shelf's seven embedding models do | |
| not agree. This one pools **mean** and | |
| **normalises**, read from | |
| `1_Pooling/config.json` and `modules.json` rather than inferred from the family name. | |
| Getting it wrong does not throw; it returns vectors that look fine and rank wrong. | |
| ## The prefix is not in the graph | |
| This model is trained with `query: ` in front of the text and expects it at | |
| inference. That happens before tokenisation, so the `.pte` never sees it as anything | |
| but tokens β and leaving it out does not throw. It returns a plausible vector that | |
| retrieves worse. | |
| ## Verification | |
| | build | file | size (MB) | Mac ms* | backend takes | worst cosine vs eager | retrieval budget | | |
| |---|---|---|---|---|---|---| | |
| | fp32 | `embed_multilingual_e5_base_xnnpack_fp32.pte` | 1110.0 | 32.5 | 77.3% | 1.000000 | 0% | | |
| | fp16 | `embed_multilingual_e5_base_xnnpack_fp16.pte` | 555.2 | 53.1 | 66.7% | 0.999999 | 6% | | |
| | Core ML (fp16, iOS) | `embed_multilingual_e5_base_coreml_all.pte` | 555.7 | 6.8 | 100.0% | 0.999994 | 19% | | |
| \*Mac arm64, one 256-token sequence, **fastest of five medians of ten** β a reference | |
| point for relative cost, not a device number. The host shares its cores with other work, | |
| and a single median does not survive that: the same eager model here measured 19.6 ms and | |
| 182.8 ms twenty minutes apart. Contention only ever adds time, so the fastest repetition is | |
| the one that means something. Torch eager fp32, measured the same way, is | |
| 34.7 ms. | |
| Cosine is measured against the model run in eager through its own pooling, over eight | |
| sentences. The last column is the one that decides: rank those eight against each | |
| other, and ask whether this build's score error is smaller than the gap between the | |
| document a query retrieves and the runner-up. Every shipped build keeps all eight | |
| top-1 results. | |
| ## The attention is eager, and that is the faster export | |
| `F.scaled_dot_product_attention` does not survive export as one operation. The edge | |
| dialect lowers it through `_safe_softmax`, whose guard against a row with no unmasked key | |
| at all leaves **11 operations XNNPACK cannot take, in every attention | |
| block** β `scalar_tensor`, `where`, `mul.Scalar`, `logical_not`, `eq`, `full_like`, `any.dim`. Each one cuts the subgraph in two. | |
| The switch is `attn_implementation="eager"`: transformers then builds the mask | |
| itself, as `torch.finfo(dtype).min`, instead of handing `F.sdpa` a **boolean** mask | |
| for PyTorch to fill with `-inf`. | |
| The guard is emitted whether or not it can ever fire, and here it cannot: it triggers only | |
| on `-inf`, and this arm never produces one. So the two differ only about rows that have no | |
| unmasked key at all β sdpa zeroes them, this one gives them a uniform row β and those are | |
| padding rows, which the pooling discards and which every real query row masks out anyway. | |
| Measured with all but eight positions masked, as adversarial as this shape gets, the two | |
| graphs agree to 1.4e-07. | |
| XNNPACK fp32 goes from **61.9% to 77.3%** delegated. | |
| ## Not shipped: int8 | |
| `embed_multilingual_e5_base_xnnpack_int8.pte` is **855.5 MB** against fp16's 555.2 MB. Dynamic int8 quantises the | |
| linear weights and leaves the token embedding table in fp32, and here that table is | |
| 768 MB of the 1110.0 MB model β **69%**. The size a build comes out at is | |
| `0.5 + 1.5 x (table share)` times the fp16 build; at 69% that is | |
| 1.54, so there was never a smaller file to be had. | |
| It is withheld on the number that decides. Ranking the eight test sentences against | |
| each other, this build moves a pair score by at most **0.0057** while the | |
| closest fp32 decision β the gap between the document a query retrieves and the | |
| runner-up β is **0.0022**. That is **255%** of | |
| the room available, against a bar of 50%. | |
| Correlation reads 0.997142 for this build, which no correlation gate | |
| would stop. | |
| torch.export -> to_edge_transform_and_lower(partitioner) -> .pte | |
| (conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models)) | |