DylanCouzon commited on
Commit
0e9cd89
·
verified ·
1 Parent(s): 6d58742

Update model card for the research preview: one FastEmbed branch for the whole family, native fp32 query vectors, canonical clean-4 results with all six beside them.

Browse files
Files changed (1) hide show
  1. README.md +119 -184
README.md CHANGED
@@ -9,269 +9,204 @@ tags:
9
  - retrieval
10
  - asymmetric-dual-encoder
11
  - edge
 
12
  base_model: NovaSearch/stella_en_400M_v5
13
  pipeline_tag: feature-extraction
14
  ---
15
 
16
  # constella-zero
17
 
18
- The query side of an **asymmetric dual encoder**: documents are indexed once, in the cloud, by a
19
- large frozen encoder; queries are encoded on the device by **a lookup table**.
 
 
 
 
20
 
21
- There is no transformer here. The model is 30,522 × 1024 int8 rows and one pooling rule —
22
- encoding a query is a gather and a weighted sum. The query asset is **31.8 MB**, and the reference
23
- implementation encodes a query end to end, tokenization included, in **0.38 ms** on one CPU core
24
- (the ONNX graph alone runs an 8-token query in 0.047 ms — see [Costs](#costs)).
25
-
26
- It was distilled from [`stella_en_400M_v5`](https://huggingface.co/NovaSearch/stella_en_400M_v5)
27
- so that its output lands in that model's document space. The matching document encoder is
28
- published as [`stella-en-400M-v5-doc-onnx`](https://huggingface.co/DylanCouzon/stella-en-400M-v5-doc-onnx);
29
- the two are only meaningful together.
30
-
31
- *constella = constellation + stella: navigate by fixed stars, no engine.*
32
-
33
- > **Research preview.** It is a bag of tokens and behaves like one. Read
34
- > [Results](#results) and [Limits](#limits) first.
35
 
36
  ## Usage
37
 
38
- The snippets in this section run in order, sharing state.
 
 
 
39
 
40
  ```python
 
41
  from fastembed import TextEmbedding
42
 
43
  NAME = "DylanCouzon/constella-zero"
44
  query_model = TextEmbedding(NAME)
45
- q = next(iter(query_model.embed(["how do mrna vaccines work?"]))) # (1024,), L2-normalized
 
46
  ```
47
 
48
- Not in a FastEmbed release yet. Until it is:
49
-
50
- pip install "fastembed @ git+https://github.com/Dylancouzon/fastembed@add-constella-models"
51
-
52
- FastEmbed fetches only `model.onnx` and the tokenizer — about 31 MB, not the whole repo. Pooling
53
- and L2 normalization happen inside the graph.
54
 
55
  ### The document side
56
 
57
  ```python
58
  DOC_NAME = "DylanCouzon/stella-en-400M-v5-doc-onnx"
59
-
60
- doc_model = TextEmbedding(DOC_NAME) # 1.75 GB, runs in the cloud, once per document
61
  docs = [
62
- "mRNA vaccines deliver a strand of messenger RNA encoding a viral antigen.",
63
  "The Treaty of Westphalia ended the Thirty Years' War in 1648.",
64
  ]
65
- D = list(doc_model.embed(docs))
 
66
  ```
67
 
68
- That asymmetry is the point: `doc_model` is a 400M-parameter transformer that runs once per
69
- document. `query_model` runs on every query, on the device, and costs almost nothing.
70
 
71
  ### With Qdrant
72
 
73
  ```python
74
  from qdrant_client import QdrantClient, models
75
 
76
- client = QdrantClient(":memory:") # or your cluster
77
- client.create_collection("docs", vectors_config=models.VectorParams(
78
- size=1024, distance=models.Distance.COSINE))
79
- client.upsert("docs", points=[
80
- models.PointStruct(id=i, vector=D[i].tolist(), payload={"text": t})
81
- for i, t in enumerate(docs)])
82
-
83
- hits = client.query_points("docs", query=q.tolist(), limit=5).points
 
 
 
 
84
  print(hits[0].payload["text"])
85
  ```
86
 
87
- Qdrant implements cosine as a dot product — it normalizes on upsert and compares with dot — so
88
- `COSINE` costs the same as `DOT` here without assuming the caller preserved unit norm.
89
-
90
- The table itself can also live in Qdrant, as a retrieve-by-id collection of one point per vocab
91
- row (`hnsw_config=models.HnswConfigDiff(m=0)` — indexing it is pure waste), so the query path holds
92
- no model weights at all.
93
-
94
  ### Without FastEmbed
95
 
96
- `zero_encoder.py` is the reference implementation — 93 lines, `numpy` and `tokenizers`, no torch.
97
- This downloads the whole repo, not just the 31 MB graph.
98
 
99
  ```python
100
  from huggingface_hub import snapshot_download
101
- import sys, numpy as np
102
 
103
  d = snapshot_download("DylanCouzon/constella-zero")
104
  sys.path.insert(0, d)
105
  from zero_encoder import ZeroQueryEncoder
106
 
107
- enc = ZeroQueryEncoder(d, variant="int8") # or "fp16"
108
- q_np = enc.encode(["how do mrna vaccines work?"]) # (1, 1024), L2-normalized
109
- assert np.abs(q_np[0] - q).max() < 1e-5 # the vector FastEmbed just produced
110
  ```
111
 
112
- ## How it works
 
113
 
114
- Tokenize (WordPiece, special tokens on, truncate at 512, no padding, no prefix). A token appearing
115
- `c` times carries **total weight `sqrt(c)`** — repetition saturates. Sum the rows, divide by the
116
- weight sum, L2-normalize. An empty or near-zero-norm bag falls back to the normalized `[CLS]` row
117
- (id 101). Per-token learned weights are folded into the rows, so the artifact is self-contained.
118
 
119
- Because pooling is not a masked mean, it is done inside the ONNX graph rather than by the caller.
120
- `config.json` carries the rule and its fingerprint (`adb24fb2e8cad66f`).
 
 
 
121
 
122
- `int8` is the variant every number below was measured on; it is loss-free against `fp16` to within
123
- 0.00013 nDCG@10.
124
 
125
  ## Files
126
 
127
- You need exactly one of these three.
128
-
129
- | file | for | size |
130
- |---|---|---|
131
- | `model.onnx` | FastEmbed, or any ONNX runtime — pooled and normalized, `(b, 1024)` | 31 MB |
132
- | `model_tokens.onnx` | pipelines that insist on pooling themselves, `(b, s, 1024)` | 31 MB |
133
- | `model.npz` | the numpy reference path | 94 MB |
134
-
135
- Both graphs are opset 17, standard operators only, carrying the table as an int8 initializer with a
136
- per-row fp32 scale dequantized in-graph.
137
 
138
- The bundled tokenizer files are stella's, with `model_max_length`/`max_length` set to **512** and
139
- `padding` to **null** — the rule the document index was built with. stella ships 32768/8000 and
140
- fixed-512 padding, which any loader honouring those fields would otherwise apply.
141
- `config.json` records the originals under `tokenizer_deviation_from_teacher`.
142
 
143
  ## Results
144
 
145
- nDCG@10 on six BEIR datasets, exact search so ANN recall is not a confound. Measured once, on the
146
- table shipped here (sha `a7007b1a…`).
 
147
 
148
- | system | arguana | fiqa | nfcorpus | scidocs | scifact | trec-covid | **average** |
149
- |---|---|---|---|---|---|---|---|
150
- | **constella-zero (int8)** | 0.5916 | 0.3728 | 0.3124 | 0.1677 | 0.6101 | 0.5490 | **0.4339** |
151
- | **+ BM25, Qdrant `Fusion.DBSF`, prefetch 100** | 0.5800 | 0.3872 | 0.3442 | 0.1850 | 0.7173 | 0.7184 | **0.4887** |
152
- | + BM25, convex fusion (not runnable in Qdrant) | 0.5975 | 0.4026 | 0.3497 | 0.1881 | 0.7068 | 0.7018 | **0.4911** |
153
- | BM25 alone | 0.4878 | 0.2532 | 0.3180 | 0.1565 | 0.6791 | 0.6099 | 0.4174 |
154
- | the teacher, used on both sides | 0.6369 | 0.5536 | 0.4134 | 0.2395 | 0.7796 | 0.8234 | 0.5744 |
155
 
156
- A lookup table retains **75.5%** of the teacher's quality (0.4339 / 0.5744), with a query side
157
- that does no matrix multiplication at all.
158
 
159
- ### Fusing with BM25 in Qdrant
 
 
 
160
 
161
- **The recommended fused system is `Fusion.DBSF` with a prefetch limit of 100** — the row in bold
162
- above. DBSF has **no fitted fusion weights**; the prefetch limit of 100 was chosen from where DBSF
163
- saturates on our development set, plus a deployability criterion, so the configuration is
164
- development-informed even though the operator itself fits nothing.
165
 
166
- Fusion needs **named** vectors, so hybrid search gets its own collection:
 
 
 
167
 
168
- ```python
169
- # The sparse side is whatever lexical model you use -- FastEmbed's `Qdrant/bm25`, or your own.
170
- # Placeholder sparse vectors here, so this snippet runs with no extra download.
171
- client.create_collection(
172
- "hybrid",
173
- vectors_config={"dense": models.VectorParams(size=1024, distance=models.Distance.COSINE)},
174
- sparse_vectors_config={"bm25": models.SparseVectorParams()},
175
- )
176
- client.upsert("hybrid", points=[
177
- models.PointStruct(
178
- id=i,
179
- vector={"dense": D[i].tolist(),
180
- "bm25": models.SparseVector(indices=[i], values=[1.0])},
181
- payload={"text": t})
182
- for i, t in enumerate(docs)])
183
-
184
- hits = client.query_points(
185
- "hybrid",
186
- prefetch=[
187
- models.Prefetch(query=q.tolist(), using="dense", limit=100),
188
- models.Prefetch(query=models.SparseVector(indices=[0], values=[1.0]),
189
- using="bm25", limit=100),
190
- ],
191
- query=models.FusionQuery(fusion=models.Fusion.DBSF),
192
- limit=10,
193
- ).points
194
- print(hits[0].payload["text"])
195
- ```
196
 
197
- **On the four datasets with no disclosed teacher overlap** (see Limits), DBSF at prefetch 100 scores
198
- **0.4912** against convex fusion's 0.4866; across all six, 0.4887 vs 0.4911. Both differences are
199
- inside the ~0.005 band we treat as noise, and we computed no confidence interval for them, so read
200
- this as **no measured quality difference in either direction** — not as DBSF being better. The
201
- reason to prefer it is that it *runs in the product*, needs no 1000-deep prefetch, and removes a
202
- tuned weight from the system.
203
-
204
- The `convex fusion` row is retained for continuity: it was the operator of record when this model
205
- was released. It is `0.8 × dense + 0.2 × BM25`, each channel divided by its per-query maximum, at
206
- prefetch depth 1000 — **Qdrant does not implement it**, and a 1000-deep prefetch to return 10
207
- results is not a realistic configuration.
208
-
209
- `Fusion.RRF` is the weaker choice. We swept it fairly — `k` from 1 to 101 in Qdrant's units (best
210
- `k=3`), and 24 weighted configurations (best `k=2, weights=[2, 1]`) — and its best point lands
211
- below DBSF on our development set. An earlier version of this card said only that RRF "will not
212
- reproduce" the fused row; that was true, but rested on an unweighted, badly-ranged comparison,
213
- which has since been redone.
214
-
215
- **Caveats.** Numbers use `bm25s` (lucene defaults), not Qdrant's own BM25, which has a fixed
216
- `avg_len` and its own tokenizer; DBSF normalises over the returned scores, so a different lexical
217
- implementation shifts its inputs.
218
-
219
- Our evaluation excludes each query's own document *before* truncating to 100, so the numbers
220
- describe a prefetch with a **self-exclusion filter** (`must_not` on the point id). Without one, a
221
- plain `limit: 100` spends a slot on the self-match. This matters only where queries are also
222
- documents — ArguAna (1,298 of 1,406 queries) and FiQA (55); the other four datasets have none, so
223
- the clean-4 figures are unaffected either way.
224
 
225
  ## Limits
226
 
227
- - **Teacher contamination.** stella discloses **ArguAna** and **FiQA** in its training data —
228
- two of the six above, and ArguAna is its second-highest score. On the four sets with no
229
- disclosed overlap it averages **0.4098 against BM25's 0.4409** — below BM25. Weight the
230
- average accordingly.
231
- - **It is a bag of tokens.** Word order, negation and syntax are not represented: "dog bites man"
232
- and "man bites dog" give the same vector.
233
- - **Out of domain it drops.** Training was Wikipedia- and e-commerce-shaped, and the six sets
234
- above are further from that than the data it was fitted on.
235
- - **English only**, 512 wordpieces, 30,522-token WordPiece vocab. Out-of-vocabulary terms degrade
236
- to subword rows.
237
- - **The document side is not cheap** — 2.05 GB per 1M documents at 1024-d fp16. The whole trade is
238
- on the query side.
239
 
240
  ## Costs
241
 
242
- | | |
243
- |---|---|
244
- | query asset (int8 rows + scales + tokenizer) | 31.8 MB |
245
- | `model.onnx` graph execution, batch 1, one thread, 8-token query | 0.047 ms |
246
- | `model.onnx` graph execution, batch 1, one thread, 512-token query | 1.22 ms |
247
- | `zero_encoder.py` end to end, batch 1, one CPU core, incl. tokenization | 0.38 ms |
248
- | hydration (cold load to first query) | 0.22 s |
249
- | document vectors, 1024-d fp16 / int8 | 2.05 / 1.02 GB per 1M — raw payload, before index overhead |
250
 
251
- The graph rows exclude tokenization; `zero_encoder.py`'s 0.38 ms is the end-to-end figure and the
252
- honest one to compare against another encoder. No end-to-end FastEmbed timing is published here.
 
 
 
253
 
254
- The graph derives token counts from an all-pairs comparison, so cost grows with the **square** of
255
- sequence length — 26x from an 8-token query to a 512-token one. Real queries sit at the short end
256
- (median 13 wordpieces).
257
 
258
  ## Training
259
 
260
- L2 regression of the table's pooled output onto the teacher's query embeddings, over 340,850
261
- pairs plus 220,632 query-text-only rows, from **Amazon ESCI**, **FEVER**, **HotpotQA**, **SQuAD**,
262
- **NQ-open**, **TriviaQA** and **Mr. TyDi (en)**. No MS MARCO.
263
-
264
- Attribution: NQ, SQuAD, HotpotQA, FEVER and Mr. TyDi are Wikipedia-derived and **CC BY-SA**
265
- (3.0/4.0); Amazon ESCI and TriviaQA are Apache-2.0; the teacher is MIT.
266
 
267
  ## Provenance
268
 
269
- ```
270
- run_id p35w-2m-s2500
271
- table sha256 a7007b1a6af120b976f093fd69ddcb5001996ec0b84b5864b4fd25d7af878abf
272
- teacher NovaSearch/stella_en_400M_v5 @ ffeb2b7ee715c226d4ffe5e4619f7dbb48624c20
273
- preproc prefix="" · add_special_tokens · max_length=512 · pool_mode=sqrt
274
- preproc fingerprint adb24fb2e8cad66f
 
275
  ```
276
 
277
- Published as `zero-query-encoder-v1` and renamed on 2026-09-03; the old URL redirects.
 
 
9
  - retrieval
10
  - asymmetric-dual-encoder
11
  - edge
12
+ - research-preview
13
  base_model: NovaSearch/stella_en_400M_v5
14
  pipeline_tag: feature-extraction
15
  ---
16
 
17
  # constella-zero
18
 
19
+ **constella-zero** is the smallest query encoder in an asymmetric retrieval family. Documents are
20
+ indexed once, in the cloud, with the frozen
21
+ [`stella-en-400M-v5-doc-onnx`](https://huggingface.co/DylanCouzon/stella-en-400M-v5-doc-onnx)
22
+ tower; Zero or the stronger [`constella-nano`](https://huggingface.co/DylanCouzon/constella-nano)
23
+ can then query that same 1024-dimensional index without re-encoding it. Zero is a 30,522 × 1024
24
+ int8 lookup table—not a transformer—so encoding is a gather and weighted sum.
25
 
26
+ > **Research preview.** The registered reserved-four evaluation and broad descriptive BEIR-18
27
+ > validation are pending and unspent; no result is claimed for either. The three models are
28
+ > registered on the preview branch below, not in an upstream FastEmbed release yet.
 
 
 
 
 
 
 
 
 
 
 
29
 
30
  ## Usage
31
 
32
+ ```console
33
+ pip install "fastembed @ git+https://github.com/Dylancouzon/fastembed.git@constella-research-preview"
34
+ pip install qdrant-client
35
+ ```
36
 
37
  ```python
38
+ import numpy as np
39
  from fastembed import TextEmbedding
40
 
41
  NAME = "DylanCouzon/constella-zero"
42
  query_model = TextEmbedding(NAME)
43
+ q = np.asarray(next(iter(query_model.embed(["how do mrna vaccines work?"]))))
44
+ assert q.shape == (1024,) and q.dtype == np.float32
45
  ```
46
 
47
+ FastEmbed fetches `model.onnx` and the tokenizer. Pooling and L2 normalization happen inside the
48
+ graph.
 
 
 
 
49
 
50
  ### The document side
51
 
52
  ```python
53
  DOC_NAME = "DylanCouzon/stella-en-400M-v5-doc-onnx"
54
+ doc_model = TextEmbedding(DOC_NAME)
 
55
  docs = [
56
+ "mRNA vaccines deliver messenger RNA encoding a viral antigen.",
57
  "The Treaty of Westphalia ended the Thirty Years' War in 1648.",
58
  ]
59
+ D = np.stack(list(doc_model.embed(docs)))
60
+ assert D.shape == (2, 1024) and D.dtype == np.float32
61
  ```
62
 
63
+ The document tower runs once per document; Zero runs on every query. Do not use the document
64
+ model's unprompted path as a Stella query encoder.
65
 
66
  ### With Qdrant
67
 
68
  ```python
69
  from qdrant_client import QdrantClient, models
70
 
71
+ client = QdrantClient(":memory:")
72
+ client.create_collection(
73
+ "docs", vectors_config=models.VectorParams(size=1024, distance=models.Distance.COSINE)
74
+ )
75
+ client.upsert(
76
+ "docs",
77
+ points=[
78
+ models.PointStruct(id=i, vector=D[i].tolist(), payload={"text": text})
79
+ for i, text in enumerate(docs)
80
+ ],
81
+ )
82
+ hits = client.query_points("docs", query=q.tolist(), limit=2).points
83
  print(hits[0].payload["text"])
84
  ```
85
 
 
 
 
 
 
 
 
86
  ### Without FastEmbed
87
 
88
+ `zero_encoder.py` is the NumPy/tokenizers reference path:
 
89
 
90
  ```python
91
  from huggingface_hub import snapshot_download
92
+ import sys
93
 
94
  d = snapshot_download("DylanCouzon/constella-zero")
95
  sys.path.insert(0, d)
96
  from zero_encoder import ZeroQueryEncoder
97
 
98
+ enc = ZeroQueryEncoder(d, variant="int8")
99
+ q_np = enc.encode(["how do mrna vaccines work?"])
100
+ assert np.abs(q_np[0] - q).max() < 1e-5
101
  ```
102
 
103
+ Zero and Nano share a document space, not a retrieval-quality guarantee: they have different
104
+ measured behavior, and interchangeability does not mean parity, equivalence, or a tie.
105
 
106
+ ## How it works
 
 
 
107
 
108
+ Tokenize with WordPiece, special tokens on, no prefix, and truncation at 512 tokens. A token that
109
+ appears `c` times carries total weight `sqrt(c)`, so repetition saturates. The graph sums the rows,
110
+ divides by the weight sum, and L2-normalizes; an empty or near-zero bag falls back to normalized
111
+ `[CLS]`. Learned token weights are folded into the rows. This is still a bag of tokens: word order,
112
+ negation, and syntax are not represented.
113
 
114
+ The reported model is int8, which was within 0.00013 nDCG@10 of fp16.
 
115
 
116
  ## Files
117
 
118
+ | file | purpose | size |
119
+ |---|---|---:|
120
+ | `model.onnx` | FastEmbed/ONNX Runtime; pooled normalized `(batch, 1024)` output | 31 MB |
121
+ | `model_tokens.onnx` | token-level `(batch, sequence, 1024)` output | 31 MB |
122
+ | `model.npz` | NumPy reference path | 94 MB |
 
 
 
 
 
123
 
124
+ Both graphs use opset 17 and standard operators. The int8 table is dequantized in-graph with one
125
+ fp32 scale per row. Tokenizer metadata enforces the frozen 512-token rule and dynamic padding.
 
 
126
 
127
  ## Results
128
 
129
+ The headline partition is **clean-4**: NFCorpus, SCIDOCS, SciFact, and TREC-COVID. ArguAna and
130
+ FiQA remain beside it with `†` because Stella discloses training/evaluation contact with them.
131
+ All values are exact-search nDCG@10.
132
 
133
+ | system | NFCorpus **(clean-4)** | SCIDOCS **(clean-4)** | SciFact **(clean-4)** | TREC-COVID **(clean-4)** | ArguAna† | FiQA† |
134
+ |---|---:|---:|---:|---:|---:|---:|
135
+ | constella-zero (int8) | 0.3124 | 0.1677 | 0.6101 | 0.5490 | 0.5916 | 0.3728 |
136
+ | constella-nano | 0.363080 | 0.217710 | 0.721097 | 0.787116 | 0.623296 | 0.477765 |
137
+ | BM25 | 0.3180 | 0.1565 | 0.6791 | 0.6099 | 0.4878 | 0.2532 |
138
+ | Stella teacher, symmetric | 0.4134 | 0.2395 | 0.7796 | 0.8234 | 0.6369 | 0.5536 |
 
139
 
140
+ † Stella-disclosed training/evaluation contact; excluded from clean-4.
 
141
 
142
+ Zero did **not** confirmatorily beat BM25. M7 C2 was +0.0165 across all six with raw 95% interval
143
+ [+0.0017, +0.0311], but its sign-flip p=0.0149 failed the Holm threshold of 0.0083. On clean-4,
144
+ the descriptive contrast was -0.0311 [-0.0517, -0.0109], with Zero at 0.4098 versus BM25 at
145
+ 0.4409. Superiority is **UNESTABLISHED**.
146
 
147
+ The deployable hybrid recommendation and registered operator of record are distinct:
 
 
 
148
 
149
+ | Zero + BM25 fusion | prefetch | all-six macro | clean-4 macro |
150
+ |---|---:|---:|---:|
151
+ | Qdrant DBSF | 100 | 0.4887 | 0.4912 |
152
+ | M7 convex0 (`w=0.8`) | 1000 | 0.4911 | 0.4866 |
153
 
154
+ Both rows use `bm25s` with Lucene defaults as the lexical side, not Qdrant's own BM25, which has a
155
+ fixed `avg_len` and its own tokenizer — DBSF normalises over returned scores, so a different
156
+ lexical implementation shifts its inputs. For the DBSF row, each query's own document is excluded
157
+ *before* the prefetch is truncated to 100 (`must_not` on the point id); without that filter a
158
+ plain `limit: 100` spends a slot on the self-match. That affects only ArguAna (1,298 of 1,406
159
+ queries) and FiQA (55), so the clean-4 figures are unchanged either way.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
160
 
161
+ Use Qdrant DBSF at prefetch 100 in deployments. M7's convex0 is the registered operator of record,
162
+ but Qdrant does not implement it. No confidence interval compared these observations, so neither
163
+ superiority nor equivalence is established. M7 C3 likewise did not establish fusion superiority
164
+ over OpenSearch: +0.0043, raw 95% interval [-0.0063, +0.0151], p=0.219.
165
+
166
+ Full per-dataset fusion rows, registered contrasts, and source traces are in the
167
+ [M21 benchmark ledger](https://github.com/Dylancouzon/asymmetric-dual-encoders/blob/main/m21/BENCHMARKS.md).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
168
 
169
  ## Limits
170
 
171
+ - Reserved-four and BEIR-18 evidence is pending and unspent; the six datasets do not establish
172
+ broad domain coverage.
173
+ - Stella contact with ArguAna and FiQA makes all-six results secondary to clean-4.
174
+ - Zero is an English-only bag of tokens and truncates beyond 512 wordpieces.
175
+ - The document side is not cheap: the 400M-parameter Stella tower still runs at indexing time.
 
 
 
 
 
 
 
176
 
177
  ## Costs
178
 
179
+ These are synthetic query latencies, not workload estimates. The common three-model protocol used
180
+ three fresh processes per model, batch one, four CPU threads; hydration includes imports,
181
+ verification, and load but excludes interpreter startup; first inference is separate; warm timing
182
+ uses five warmups and twenty 20-word samples; the OS disk cache was not flushed.
 
 
 
 
183
 
184
+ | model | hydration | first query | warm 20-word p50 | peak RSS | model assets |
185
+ |---|---:|---:|---:|---:|---:|
186
+ | constella-zero | 0.2618 s | 0.3529 ms | 0.1119 ms | 275.4 MiB | 90.1 MiB |
187
+ | bge-small | 0.6726 s | 8.2401 ms | 6.8400 ms | 291.0 MiB | 127.6 MiB |
188
+ | constella-nano | 0.6907 s | 7.6685 ms | 7.2511 ms | 280.9 MiB | 132.3 MiB |
189
 
190
+ This measures query encoders only, not retrieval, ANN, or end-to-end system latency. The asset
191
+ column follows the common protocol; `model.onnx` itself is the 31 MB query graph listed above.
 
192
 
193
  ## Training
194
 
195
+ The table was trained by L2 regression against Stella query embeddings over 340,850 pairs plus
196
+ 220,632 query-only rows from Amazon ESCI, FEVER, HotpotQA, SQuAD, NQ Open, TriviaQA, and Mr. TyDi
197
+ (English); MS MARCO was excluded. Wikipedia-derived sources retain CC BY-SA attribution; ESCI and
198
+ TriviaQA are Apache-2.0.
 
 
199
 
200
  ## Provenance
201
 
202
+ ```text
203
+ run_id p35w-2m-s2500
204
+ table a7007b1a6af120b976f093fd69ddcb5001996ec0b84b5864b4fd25d7af878abf
205
+ teacher NovaSearch/stella_en_400M_v5 @ ffeb2b7ee715c226d4ffe5e4619f7dbb48624c20
206
+ comparator BAAI/bge-small-en-v1.5 @ 5c38ec7c405ec4b44b94cc5a9bb96e735b38267a
207
+ preproc prefix="" · special tokens · max_length=512 · pool_mode=sqrt
208
+ fingerprint adb24fb2e8cad66f
209
  ```
210
 
211
+ The weights are MIT licensed. The pinned Stella teacher/document tower and bge-small comparator
212
+ are also MIT and are attributed above.