zero v1 — M7 lookup table, run p35w-2m-s2500
Browse files- README.md +41 -1
- model.onnx +3 -0
- model_tokens.onnx +3 -0
README.md
CHANGED
|
@@ -44,6 +44,37 @@ enc = ZeroQueryEncoder(d, variant="int8") # or "fp16"
|
|
| 44 |
q = enc.encode(["how do mrna vaccines work?"]) # (1, 1024), L2-normalized
|
| 45 |
```
|
| 46 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 47 |
Documents are encoded by the frozen teacher — **pin the revision**, the table is only valid
|
| 48 |
against this exact document space:
|
| 49 |
|
|
@@ -121,6 +152,8 @@ fingerprint (`adb24fb2e8cad66f`).
|
|
| 121 |
| file | what |
|
| 122 |
|---|---|
|
| 123 |
| `model.npz` | `rows_int8` (30522×1024) + `int8_scale`, and `rows_fp16` for reference |
|
|
|
|
|
|
|
| 124 |
| `config.json` | the frozen preprocessing rule, teacher pin, document-encoder spec, shas |
|
| 125 |
| `zero_encoder.py` | the whole query path — numpy + tokenizers, no torch |
|
| 126 |
| `tokenizer.json`, `vocab.txt`, … | stella's WordPiece tokenizer, copied at the pinned revision — with the two edits below |
|
|
@@ -194,10 +227,16 @@ while its query side remains a table lookup plus token counts.
|
|
| 194 |
|---|---|
|
| 195 |
| query asset (int8 rows + scales + tokenizer) | **31.8 MB** |
|
| 196 |
| query encode, batch 1, one CPU core | **0.38 ms** (`zero_encoder.py` measures ~0.07 ms) |
|
|
|
|
|
|
|
| 197 |
| hydration (cold load to first query) | **0.22 s** |
|
| 198 |
| document index, 1024-d fp16 | 2.05 GB per 1M documents |
|
| 199 |
| document index, 1024-d int8 | 1.02 GB per 1M documents |
|
| 200 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 201 |
For reference at the document side: LightRetriever 3.07, OpenSearch sparse 1.40,
|
| 202 |
bge-small 0.77 GB/1M. `zero`'s document index is not cheap — the trade is all on the query side.
|
| 203 |
|
|
@@ -221,7 +260,8 @@ Amazon ESCI is Apache-2.0. The teacher, `NovaSearch/stella_en_400M_v5`, is MIT.
|
|
| 221 |
| | |
|
| 222 |
|---|---|
|
| 223 |
| first published | 2026-09-03 — the frozen M7 bundle, with stella's tokenizer files copied verbatim |
|
| 224 |
-
|
|
|
|
|
| 225 |
|
| 226 |
`model.npz` is byte-identical across both (sha `a7007b1a…`) and the reference encoder's output is
|
| 227 |
unchanged, so **no published number differs between revisions**. Pass `revision=` to
|
|
|
|
| 44 |
q = enc.encode(["how do mrna vaccines work?"]) # (1, 1024), L2-normalized
|
| 45 |
```
|
| 46 |
|
| 47 |
+
### ONNX
|
| 48 |
+
|
| 49 |
+
The same query path is also an ONNX graph, so you can serve it without numpy or `tokenizers`
|
| 50 |
+
Python — no transformer, still a gather and a sum. `model.onnx` returns the pooled, normalized
|
| 51 |
+
vector; `model_tokens.onnx` returns one weighted row per token, for pipelines that do their own
|
| 52 |
+
masked-mean pooling (fastembed's, for instance).
|
| 53 |
+
|
| 54 |
+
```python
|
| 55 |
+
# pip install onnxruntime tokenizers
|
| 56 |
+
import numpy as np, onnxruntime as ort
|
| 57 |
+
from tokenizers import Tokenizer
|
| 58 |
+
|
| 59 |
+
tok = Tokenizer.from_file(f"{d}/tokenizer.json") # padding/truncation already correct
|
| 60 |
+
sess = ort.InferenceSession(f"{d}/model.onnx", providers=["CPUExecutionProvider"])
|
| 61 |
+
|
| 62 |
+
encs = [tok.encode(t) for t in ["how do mrna vaccines work?", "who invented the barometer"]]
|
| 63 |
+
L = max(len(e.ids) for e in encs)
|
| 64 |
+
ids = np.zeros((len(encs), L), np.int64)
|
| 65 |
+
attn = np.zeros((len(encs), L), np.int64)
|
| 66 |
+
for i, e in enumerate(encs):
|
| 67 |
+
ids[i, :len(e.ids)], attn[i, :len(e.ids)] = e.ids, 1
|
| 68 |
+
|
| 69 |
+
Q = sess.run(None, {"input_ids": ids, "attention_mask": attn})[0] # (2, 1024), normalized
|
| 70 |
+
assert Q.shape == (2, 1024)
|
| 71 |
+
```
|
| 72 |
+
|
| 73 |
+
Both graphs are opset 17, standard operators only, and carry the table as an **int8 initializer
|
| 74 |
+
with a per-row fp32 scale** dequantized in-graph — so each is ~31 MB rather than the 125 MB fp32
|
| 75 |
+
rows would cost. **You need exactly one of `model.npz`, `model.onnx` or `model_tokens.onnx`**, not
|
| 76 |
+
all three; the repo ships all of them so you can pick your runtime.
|
| 77 |
+
|
| 78 |
Documents are encoded by the frozen teacher — **pin the revision**, the table is only valid
|
| 79 |
against this exact document space:
|
| 80 |
|
|
|
|
| 152 |
| file | what |
|
| 153 |
|---|---|
|
| 154 |
| `model.npz` | `rows_int8` (30522×1024) + `int8_scale`, and `rows_fp16` for reference |
|
| 155 |
+
| `model.onnx` | the whole query path as one opset-17 graph → `(b, 1024)` pooled + normalized |
|
| 156 |
+
| `model_tokens.onnx` | the same, emitting `(b, s, 1024)` per-token weighted rows for pipelines that pool themselves |
|
| 157 |
| `config.json` | the frozen preprocessing rule, teacher pin, document-encoder spec, shas |
|
| 158 |
| `zero_encoder.py` | the whole query path — numpy + tokenizers, no torch |
|
| 159 |
| `tokenizer.json`, `vocab.txt`, … | stella's WordPiece tokenizer, copied at the pinned revision — with the two edits below |
|
|
|
|
| 227 |
|---|---|
|
| 228 |
| query asset (int8 rows + scales + tokenizer) | **31.8 MB** |
|
| 229 |
| query encode, batch 1, one CPU core | **0.38 ms** (`zero_encoder.py` measures ~0.07 ms) |
|
| 230 |
+
| `model.onnx`, batch 1, one thread, 8-token query | **0.047 ms** |
|
| 231 |
+
| `model.onnx`, batch 1, one thread, 512-token query | **1.22 ms** |
|
| 232 |
| hydration (cold load to first query) | **0.22 s** |
|
| 233 |
| document index, 1024-d fp16 | 2.05 GB per 1M documents |
|
| 234 |
| document index, 1024-d int8 | 1.02 GB per 1M documents |
|
| 235 |
|
| 236 |
+
The ONNX graph derives token counts from an all-pairs comparison, so its cost grows with the
|
| 237 |
+
SQUARE of the sequence length — 26x from an 8-token query to a 512-token one. Real queries sit at
|
| 238 |
+
the short end (the dev set's median is 13 wordpieces), but a long one is not free.
|
| 239 |
+
|
| 240 |
For reference at the document side: LightRetriever 3.07, OpenSearch sparse 1.40,
|
| 241 |
bge-small 0.77 GB/1M. `zero`'s document index is not cheap — the trade is all on the query side.
|
| 242 |
|
|
|
|
| 260 |
| | |
|
| 261 |
|---|---|
|
| 262 |
| first published | 2026-09-03 — the frozen M7 bundle, with stella's tokenizer files copied verbatim |
|
| 263 |
+
| ONNX added | 2026-09-03 — `model.onnx` and `model_tokens.onnx`; the `.npz` and its numbers unchanged |
|
| 264 |
+
| tokenizer fixed | 2026-09-03, commit `1aa60418` — `tokenizer_config.json` `model_max_length`/`max_length` 32768/8000 → 512, `tokenizer.json` `padding` `Fixed(512)` → `null`, and one broken snippet in this card fixed |
|
| 265 |
|
| 266 |
`model.npz` is byte-identical across both (sha `a7007b1a…`) and the reference encoder's output is
|
| 267 |
unchanged, so **no published number differs between revisions**. Pass `revision=` to
|
model.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6c8a9d0753330cb5291b4df005ecfaa04a27b587f6af6f0986778dc43fb9ec0a
|
| 3 |
+
size 31382036
|
model_tokens.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:457e73ec07c958982fcc7a276e4c218d6ab70b8a2c55452e38a318c27dcfb5f0
|
| 3 |
+
size 31377524
|