DylanCouzon commited on
Commit
fb8e5c5
·
verified ·
1 Parent(s): 1aa6041

zero v1 — M7 lookup table, run p35w-2m-s2500

Browse files
Files changed (3) hide show
  1. README.md +41 -1
  2. model.onnx +3 -0
  3. model_tokens.onnx +3 -0
README.md CHANGED
@@ -44,6 +44,37 @@ enc = ZeroQueryEncoder(d, variant="int8") # or "fp16"
44
  q = enc.encode(["how do mrna vaccines work?"]) # (1, 1024), L2-normalized
45
  ```
46
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
47
  Documents are encoded by the frozen teacher — **pin the revision**, the table is only valid
48
  against this exact document space:
49
 
@@ -121,6 +152,8 @@ fingerprint (`adb24fb2e8cad66f`).
121
  | file | what |
122
  |---|---|
123
  | `model.npz` | `rows_int8` (30522×1024) + `int8_scale`, and `rows_fp16` for reference |
 
 
124
  | `config.json` | the frozen preprocessing rule, teacher pin, document-encoder spec, shas |
125
  | `zero_encoder.py` | the whole query path — numpy + tokenizers, no torch |
126
  | `tokenizer.json`, `vocab.txt`, … | stella's WordPiece tokenizer, copied at the pinned revision — with the two edits below |
@@ -194,10 +227,16 @@ while its query side remains a table lookup plus token counts.
194
  |---|---|
195
  | query asset (int8 rows + scales + tokenizer) | **31.8 MB** |
196
  | query encode, batch 1, one CPU core | **0.38 ms** (`zero_encoder.py` measures ~0.07 ms) |
 
 
197
  | hydration (cold load to first query) | **0.22 s** |
198
  | document index, 1024-d fp16 | 2.05 GB per 1M documents |
199
  | document index, 1024-d int8 | 1.02 GB per 1M documents |
200
 
 
 
 
 
201
  For reference at the document side: LightRetriever 3.07, OpenSearch sparse 1.40,
202
  bge-small 0.77 GB/1M. `zero`'s document index is not cheap — the trade is all on the query side.
203
 
@@ -221,7 +260,8 @@ Amazon ESCI is Apache-2.0. The teacher, `NovaSearch/stella_en_400M_v5`, is MIT.
221
  | | |
222
  |---|---|
223
  | first published | 2026-09-03 — the frozen M7 bundle, with stella's tokenizer files copied verbatim |
224
- | this revision | 2026-09-03, commit `f5479985` — `tokenizer_config.json` `model_max_length`/`max_length` 32768/8000 → 512, `tokenizer.json` `padding` `Fixed(512)` → `null`, and one broken snippet in this card fixed |
 
225
 
226
  `model.npz` is byte-identical across both (sha `a7007b1a…`) and the reference encoder's output is
227
  unchanged, so **no published number differs between revisions**. Pass `revision=` to
 
44
  q = enc.encode(["how do mrna vaccines work?"]) # (1, 1024), L2-normalized
45
  ```
46
 
47
+ ### ONNX
48
+
49
+ The same query path is also an ONNX graph, so you can serve it without numpy or `tokenizers`
50
+ Python — no transformer, still a gather and a sum. `model.onnx` returns the pooled, normalized
51
+ vector; `model_tokens.onnx` returns one weighted row per token, for pipelines that do their own
52
+ masked-mean pooling (fastembed's, for instance).
53
+
54
+ ```python
55
+ # pip install onnxruntime tokenizers
56
+ import numpy as np, onnxruntime as ort
57
+ from tokenizers import Tokenizer
58
+
59
+ tok = Tokenizer.from_file(f"{d}/tokenizer.json") # padding/truncation already correct
60
+ sess = ort.InferenceSession(f"{d}/model.onnx", providers=["CPUExecutionProvider"])
61
+
62
+ encs = [tok.encode(t) for t in ["how do mrna vaccines work?", "who invented the barometer"]]
63
+ L = max(len(e.ids) for e in encs)
64
+ ids = np.zeros((len(encs), L), np.int64)
65
+ attn = np.zeros((len(encs), L), np.int64)
66
+ for i, e in enumerate(encs):
67
+ ids[i, :len(e.ids)], attn[i, :len(e.ids)] = e.ids, 1
68
+
69
+ Q = sess.run(None, {"input_ids": ids, "attention_mask": attn})[0] # (2, 1024), normalized
70
+ assert Q.shape == (2, 1024)
71
+ ```
72
+
73
+ Both graphs are opset 17, standard operators only, and carry the table as an **int8 initializer
74
+ with a per-row fp32 scale** dequantized in-graph — so each is ~31 MB rather than the 125 MB fp32
75
+ rows would cost. **You need exactly one of `model.npz`, `model.onnx` or `model_tokens.onnx`**, not
76
+ all three; the repo ships all of them so you can pick your runtime.
77
+
78
  Documents are encoded by the frozen teacher — **pin the revision**, the table is only valid
79
  against this exact document space:
80
 
 
152
  | file | what |
153
  |---|---|
154
  | `model.npz` | `rows_int8` (30522×1024) + `int8_scale`, and `rows_fp16` for reference |
155
+ | `model.onnx` | the whole query path as one opset-17 graph → `(b, 1024)` pooled + normalized |
156
+ | `model_tokens.onnx` | the same, emitting `(b, s, 1024)` per-token weighted rows for pipelines that pool themselves |
157
  | `config.json` | the frozen preprocessing rule, teacher pin, document-encoder spec, shas |
158
  | `zero_encoder.py` | the whole query path — numpy + tokenizers, no torch |
159
  | `tokenizer.json`, `vocab.txt`, … | stella's WordPiece tokenizer, copied at the pinned revision — with the two edits below |
 
227
  |---|---|
228
  | query asset (int8 rows + scales + tokenizer) | **31.8 MB** |
229
  | query encode, batch 1, one CPU core | **0.38 ms** (`zero_encoder.py` measures ~0.07 ms) |
230
+ | `model.onnx`, batch 1, one thread, 8-token query | **0.047 ms** |
231
+ | `model.onnx`, batch 1, one thread, 512-token query | **1.22 ms** |
232
  | hydration (cold load to first query) | **0.22 s** |
233
  | document index, 1024-d fp16 | 2.05 GB per 1M documents |
234
  | document index, 1024-d int8 | 1.02 GB per 1M documents |
235
 
236
+ The ONNX graph derives token counts from an all-pairs comparison, so its cost grows with the
237
+ SQUARE of the sequence length — 26x from an 8-token query to a 512-token one. Real queries sit at
238
+ the short end (the dev set's median is 13 wordpieces), but a long one is not free.
239
+
240
  For reference at the document side: LightRetriever 3.07, OpenSearch sparse 1.40,
241
  bge-small 0.77 GB/1M. `zero`'s document index is not cheap — the trade is all on the query side.
242
 
 
260
  | | |
261
  |---|---|
262
  | first published | 2026-09-03 — the frozen M7 bundle, with stella's tokenizer files copied verbatim |
263
+ | ONNX added | 2026-09-03 — `model.onnx` and `model_tokens.onnx`; the `.npz` and its numbers unchanged |
264
+ | tokenizer fixed | 2026-09-03, commit `1aa60418` — `tokenizer_config.json` `model_max_length`/`max_length` 32768/8000 → 512, `tokenizer.json` `padding` `Fixed(512)` → `null`, and one broken snippet in this card fixed |
265
 
266
  `model.npz` is byte-identical across both (sha `a7007b1a…`) and the reference encoder's output is
267
  unchanged, so **no published number differs between revisions**. Pass `revision=` to
model.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6c8a9d0753330cb5291b4df005ecfaa04a27b587f6af6f0986778dc43fb9ec0a
3
+ size 31382036
model_tokens.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:457e73ec07c958982fcc7a276e4c218d6ab70b8a2c55452e38a318c27dcfb5f0
3
+ size 31377524