Update model card for the research preview: one FastEmbed branch for the whole family, native fp32 query vectors, canonical clean-4 results with all six beside them.
Browse files
README.md
CHANGED
|
@@ -9,269 +9,204 @@ tags:
|
|
| 9 |
- retrieval
|
| 10 |
- asymmetric-dual-encoder
|
| 11 |
- edge
|
|
|
|
| 12 |
base_model: NovaSearch/stella_en_400M_v5
|
| 13 |
pipeline_tag: feature-extraction
|
| 14 |
---
|
| 15 |
|
| 16 |
# constella-zero
|
| 17 |
|
| 18 |
-
|
| 19 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 20 |
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
(the ONNX graph alone runs an 8-token query in 0.047 ms — see [Costs](#costs)).
|
| 25 |
-
|
| 26 |
-
It was distilled from [`stella_en_400M_v5`](https://huggingface.co/NovaSearch/stella_en_400M_v5)
|
| 27 |
-
so that its output lands in that model's document space. The matching document encoder is
|
| 28 |
-
published as [`stella-en-400M-v5-doc-onnx`](https://huggingface.co/DylanCouzon/stella-en-400M-v5-doc-onnx);
|
| 29 |
-
the two are only meaningful together.
|
| 30 |
-
|
| 31 |
-
*constella = constellation + stella: navigate by fixed stars, no engine.*
|
| 32 |
-
|
| 33 |
-
> **Research preview.** It is a bag of tokens and behaves like one. Read
|
| 34 |
-
> [Results](#results) and [Limits](#limits) first.
|
| 35 |
|
| 36 |
## Usage
|
| 37 |
|
| 38 |
-
|
|
|
|
|
|
|
|
|
|
| 39 |
|
| 40 |
```python
|
|
|
|
| 41 |
from fastembed import TextEmbedding
|
| 42 |
|
| 43 |
NAME = "DylanCouzon/constella-zero"
|
| 44 |
query_model = TextEmbedding(NAME)
|
| 45 |
-
q = next(iter(query_model.embed(["how do mrna vaccines work?"])))
|
|
|
|
| 46 |
```
|
| 47 |
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
pip install "fastembed @ git+https://github.com/Dylancouzon/fastembed@add-constella-models"
|
| 51 |
-
|
| 52 |
-
FastEmbed fetches only `model.onnx` and the tokenizer — about 31 MB, not the whole repo. Pooling
|
| 53 |
-
and L2 normalization happen inside the graph.
|
| 54 |
|
| 55 |
### The document side
|
| 56 |
|
| 57 |
```python
|
| 58 |
DOC_NAME = "DylanCouzon/stella-en-400M-v5-doc-onnx"
|
| 59 |
-
|
| 60 |
-
doc_model = TextEmbedding(DOC_NAME) # 1.75 GB, runs in the cloud, once per document
|
| 61 |
docs = [
|
| 62 |
-
"mRNA vaccines deliver
|
| 63 |
"The Treaty of Westphalia ended the Thirty Years' War in 1648.",
|
| 64 |
]
|
| 65 |
-
D = list(doc_model.embed(docs))
|
|
|
|
| 66 |
```
|
| 67 |
|
| 68 |
-
|
| 69 |
-
|
| 70 |
|
| 71 |
### With Qdrant
|
| 72 |
|
| 73 |
```python
|
| 74 |
from qdrant_client import QdrantClient, models
|
| 75 |
|
| 76 |
-
client = QdrantClient(":memory:")
|
| 77 |
-
client.create_collection(
|
| 78 |
-
size=1024, distance=models.Distance.COSINE)
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 84 |
print(hits[0].payload["text"])
|
| 85 |
```
|
| 86 |
|
| 87 |
-
Qdrant implements cosine as a dot product — it normalizes on upsert and compares with dot — so
|
| 88 |
-
`COSINE` costs the same as `DOT` here without assuming the caller preserved unit norm.
|
| 89 |
-
|
| 90 |
-
The table itself can also live in Qdrant, as a retrieve-by-id collection of one point per vocab
|
| 91 |
-
row (`hnsw_config=models.HnswConfigDiff(m=0)` — indexing it is pure waste), so the query path holds
|
| 92 |
-
no model weights at all.
|
| 93 |
-
|
| 94 |
### Without FastEmbed
|
| 95 |
|
| 96 |
-
`zero_encoder.py` is the
|
| 97 |
-
This downloads the whole repo, not just the 31 MB graph.
|
| 98 |
|
| 99 |
```python
|
| 100 |
from huggingface_hub import snapshot_download
|
| 101 |
-
import sys
|
| 102 |
|
| 103 |
d = snapshot_download("DylanCouzon/constella-zero")
|
| 104 |
sys.path.insert(0, d)
|
| 105 |
from zero_encoder import ZeroQueryEncoder
|
| 106 |
|
| 107 |
-
enc = ZeroQueryEncoder(d, variant="int8")
|
| 108 |
-
q_np = enc.encode(["how do mrna vaccines work?"])
|
| 109 |
-
assert np.abs(q_np[0] - q).max() < 1e-5
|
| 110 |
```
|
| 111 |
|
| 112 |
-
|
|
|
|
| 113 |
|
| 114 |
-
|
| 115 |
-
`c` times carries **total weight `sqrt(c)`** — repetition saturates. Sum the rows, divide by the
|
| 116 |
-
weight sum, L2-normalize. An empty or near-zero-norm bag falls back to the normalized `[CLS]` row
|
| 117 |
-
(id 101). Per-token learned weights are folded into the rows, so the artifact is self-contained.
|
| 118 |
|
| 119 |
-
|
| 120 |
-
`
|
|
|
|
|
|
|
|
|
|
| 121 |
|
| 122 |
-
|
| 123 |
-
0.00013 nDCG@10.
|
| 124 |
|
| 125 |
## Files
|
| 126 |
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
|
| 130 |
-
|
|
| 131 |
-
| `model.
|
| 132 |
-
| `model_tokens.onnx` | pipelines that insist on pooling themselves, `(b, s, 1024)` | 31 MB |
|
| 133 |
-
| `model.npz` | the numpy reference path | 94 MB |
|
| 134 |
-
|
| 135 |
-
Both graphs are opset 17, standard operators only, carrying the table as an int8 initializer with a
|
| 136 |
-
per-row fp32 scale dequantized in-graph.
|
| 137 |
|
| 138 |
-
|
| 139 |
-
|
| 140 |
-
fixed-512 padding, which any loader honouring those fields would otherwise apply.
|
| 141 |
-
`config.json` records the originals under `tokenizer_deviation_from_teacher`.
|
| 142 |
|
| 143 |
## Results
|
| 144 |
|
| 145 |
-
|
| 146 |
-
|
|
|
|
| 147 |
|
| 148 |
-
| system |
|
| 149 |
-
|---|---|---|---|---|---|---|
|
| 150 |
-
|
|
| 151 |
-
|
|
| 152 |
-
|
|
| 153 |
-
|
|
| 154 |
-
| the teacher, used on both sides | 0.6369 | 0.5536 | 0.4134 | 0.2395 | 0.7796 | 0.8234 | 0.5744 |
|
| 155 |
|
| 156 |
-
|
| 157 |
-
that does no matrix multiplication at all.
|
| 158 |
|
| 159 |
-
|
|
|
|
|
|
|
|
|
|
| 160 |
|
| 161 |
-
|
| 162 |
-
above. DBSF has **no fitted fusion weights**; the prefetch limit of 100 was chosen from where DBSF
|
| 163 |
-
saturates on our development set, plus a deployability criterion, so the configuration is
|
| 164 |
-
development-informed even though the operator itself fits nothing.
|
| 165 |
|
| 166 |
-
|
|
|
|
|
|
|
|
|
|
| 167 |
|
| 168 |
-
``
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
sparse_vectors_config={"bm25": models.SparseVectorParams()},
|
| 175 |
-
)
|
| 176 |
-
client.upsert("hybrid", points=[
|
| 177 |
-
models.PointStruct(
|
| 178 |
-
id=i,
|
| 179 |
-
vector={"dense": D[i].tolist(),
|
| 180 |
-
"bm25": models.SparseVector(indices=[i], values=[1.0])},
|
| 181 |
-
payload={"text": t})
|
| 182 |
-
for i, t in enumerate(docs)])
|
| 183 |
-
|
| 184 |
-
hits = client.query_points(
|
| 185 |
-
"hybrid",
|
| 186 |
-
prefetch=[
|
| 187 |
-
models.Prefetch(query=q.tolist(), using="dense", limit=100),
|
| 188 |
-
models.Prefetch(query=models.SparseVector(indices=[0], values=[1.0]),
|
| 189 |
-
using="bm25", limit=100),
|
| 190 |
-
],
|
| 191 |
-
query=models.FusionQuery(fusion=models.Fusion.DBSF),
|
| 192 |
-
limit=10,
|
| 193 |
-
).points
|
| 194 |
-
print(hits[0].payload["text"])
|
| 195 |
-
```
|
| 196 |
|
| 197 |
-
|
| 198 |
-
|
| 199 |
-
|
| 200 |
-
|
| 201 |
-
|
| 202 |
-
|
| 203 |
-
|
| 204 |
-
The `convex fusion` row is retained for continuity: it was the operator of record when this model
|
| 205 |
-
was released. It is `0.8 × dense + 0.2 × BM25`, each channel divided by its per-query maximum, at
|
| 206 |
-
prefetch depth 1000 — **Qdrant does not implement it**, and a 1000-deep prefetch to return 10
|
| 207 |
-
results is not a realistic configuration.
|
| 208 |
-
|
| 209 |
-
`Fusion.RRF` is the weaker choice. We swept it fairly — `k` from 1 to 101 in Qdrant's units (best
|
| 210 |
-
`k=3`), and 24 weighted configurations (best `k=2, weights=[2, 1]`) — and its best point lands
|
| 211 |
-
below DBSF on our development set. An earlier version of this card said only that RRF "will not
|
| 212 |
-
reproduce" the fused row; that was true, but rested on an unweighted, badly-ranged comparison,
|
| 213 |
-
which has since been redone.
|
| 214 |
-
|
| 215 |
-
**Caveats.** Numbers use `bm25s` (lucene defaults), not Qdrant's own BM25, which has a fixed
|
| 216 |
-
`avg_len` and its own tokenizer; DBSF normalises over the returned scores, so a different lexical
|
| 217 |
-
implementation shifts its inputs.
|
| 218 |
-
|
| 219 |
-
Our evaluation excludes each query's own document *before* truncating to 100, so the numbers
|
| 220 |
-
describe a prefetch with a **self-exclusion filter** (`must_not` on the point id). Without one, a
|
| 221 |
-
plain `limit: 100` spends a slot on the self-match. This matters only where queries are also
|
| 222 |
-
documents — ArguAna (1,298 of 1,406 queries) and FiQA (55); the other four datasets have none, so
|
| 223 |
-
the clean-4 figures are unaffected either way.
|
| 224 |
|
| 225 |
## Limits
|
| 226 |
|
| 227 |
-
-
|
| 228 |
-
|
| 229 |
-
|
| 230 |
-
|
| 231 |
-
-
|
| 232 |
-
and "man bites dog" give the same vector.
|
| 233 |
-
- **Out of domain it drops.** Training was Wikipedia- and e-commerce-shaped, and the six sets
|
| 234 |
-
above are further from that than the data it was fitted on.
|
| 235 |
-
- **English only**, 512 wordpieces, 30,522-token WordPiece vocab. Out-of-vocabulary terms degrade
|
| 236 |
-
to subword rows.
|
| 237 |
-
- **The document side is not cheap** — 2.05 GB per 1M documents at 1024-d fp16. The whole trade is
|
| 238 |
-
on the query side.
|
| 239 |
|
| 240 |
## Costs
|
| 241 |
|
| 242 |
-
|
| 243 |
-
|
| 244 |
-
|
| 245 |
-
|
| 246 |
-
| `model.onnx` graph execution, batch 1, one thread, 512-token query | 1.22 ms |
|
| 247 |
-
| `zero_encoder.py` end to end, batch 1, one CPU core, incl. tokenization | 0.38 ms |
|
| 248 |
-
| hydration (cold load to first query) | 0.22 s |
|
| 249 |
-
| document vectors, 1024-d fp16 / int8 | 2.05 / 1.02 GB per 1M — raw payload, before index overhead |
|
| 250 |
|
| 251 |
-
|
| 252 |
-
|
|
|
|
|
|
|
|
|
|
| 253 |
|
| 254 |
-
|
| 255 |
-
|
| 256 |
-
(median 13 wordpieces).
|
| 257 |
|
| 258 |
## Training
|
| 259 |
|
| 260 |
-
|
| 261 |
-
|
| 262 |
-
|
| 263 |
-
|
| 264 |
-
Attribution: NQ, SQuAD, HotpotQA, FEVER and Mr. TyDi are Wikipedia-derived and **CC BY-SA**
|
| 265 |
-
(3.0/4.0); Amazon ESCI and TriviaQA are Apache-2.0; the teacher is MIT.
|
| 266 |
|
| 267 |
## Provenance
|
| 268 |
|
| 269 |
-
```
|
| 270 |
-
run_id
|
| 271 |
-
table
|
| 272 |
-
teacher
|
| 273 |
-
|
| 274 |
-
preproc
|
|
|
|
| 275 |
```
|
| 276 |
|
| 277 |
-
|
|
|
|
|
|
| 9 |
- retrieval
|
| 10 |
- asymmetric-dual-encoder
|
| 11 |
- edge
|
| 12 |
+
- research-preview
|
| 13 |
base_model: NovaSearch/stella_en_400M_v5
|
| 14 |
pipeline_tag: feature-extraction
|
| 15 |
---
|
| 16 |
|
| 17 |
# constella-zero
|
| 18 |
|
| 19 |
+
**constella-zero** is the smallest query encoder in an asymmetric retrieval family. Documents are
|
| 20 |
+
indexed once, in the cloud, with the frozen
|
| 21 |
+
[`stella-en-400M-v5-doc-onnx`](https://huggingface.co/DylanCouzon/stella-en-400M-v5-doc-onnx)
|
| 22 |
+
tower; Zero or the stronger [`constella-nano`](https://huggingface.co/DylanCouzon/constella-nano)
|
| 23 |
+
can then query that same 1024-dimensional index without re-encoding it. Zero is a 30,522 × 1024
|
| 24 |
+
int8 lookup table—not a transformer—so encoding is a gather and weighted sum.
|
| 25 |
|
| 26 |
+
> **Research preview.** The registered reserved-four evaluation and broad descriptive BEIR-18
|
| 27 |
+
> validation are pending and unspent; no result is claimed for either. The three models are
|
| 28 |
+
> registered on the preview branch below, not in an upstream FastEmbed release yet.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 29 |
|
| 30 |
## Usage
|
| 31 |
|
| 32 |
+
```console
|
| 33 |
+
pip install "fastembed @ git+https://github.com/Dylancouzon/fastembed.git@constella-research-preview"
|
| 34 |
+
pip install qdrant-client
|
| 35 |
+
```
|
| 36 |
|
| 37 |
```python
|
| 38 |
+
import numpy as np
|
| 39 |
from fastembed import TextEmbedding
|
| 40 |
|
| 41 |
NAME = "DylanCouzon/constella-zero"
|
| 42 |
query_model = TextEmbedding(NAME)
|
| 43 |
+
q = np.asarray(next(iter(query_model.embed(["how do mrna vaccines work?"]))))
|
| 44 |
+
assert q.shape == (1024,) and q.dtype == np.float32
|
| 45 |
```
|
| 46 |
|
| 47 |
+
FastEmbed fetches `model.onnx` and the tokenizer. Pooling and L2 normalization happen inside the
|
| 48 |
+
graph.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 49 |
|
| 50 |
### The document side
|
| 51 |
|
| 52 |
```python
|
| 53 |
DOC_NAME = "DylanCouzon/stella-en-400M-v5-doc-onnx"
|
| 54 |
+
doc_model = TextEmbedding(DOC_NAME)
|
|
|
|
| 55 |
docs = [
|
| 56 |
+
"mRNA vaccines deliver messenger RNA encoding a viral antigen.",
|
| 57 |
"The Treaty of Westphalia ended the Thirty Years' War in 1648.",
|
| 58 |
]
|
| 59 |
+
D = np.stack(list(doc_model.embed(docs)))
|
| 60 |
+
assert D.shape == (2, 1024) and D.dtype == np.float32
|
| 61 |
```
|
| 62 |
|
| 63 |
+
The document tower runs once per document; Zero runs on every query. Do not use the document
|
| 64 |
+
model's unprompted path as a Stella query encoder.
|
| 65 |
|
| 66 |
### With Qdrant
|
| 67 |
|
| 68 |
```python
|
| 69 |
from qdrant_client import QdrantClient, models
|
| 70 |
|
| 71 |
+
client = QdrantClient(":memory:")
|
| 72 |
+
client.create_collection(
|
| 73 |
+
"docs", vectors_config=models.VectorParams(size=1024, distance=models.Distance.COSINE)
|
| 74 |
+
)
|
| 75 |
+
client.upsert(
|
| 76 |
+
"docs",
|
| 77 |
+
points=[
|
| 78 |
+
models.PointStruct(id=i, vector=D[i].tolist(), payload={"text": text})
|
| 79 |
+
for i, text in enumerate(docs)
|
| 80 |
+
],
|
| 81 |
+
)
|
| 82 |
+
hits = client.query_points("docs", query=q.tolist(), limit=2).points
|
| 83 |
print(hits[0].payload["text"])
|
| 84 |
```
|
| 85 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 86 |
### Without FastEmbed
|
| 87 |
|
| 88 |
+
`zero_encoder.py` is the NumPy/tokenizers reference path:
|
|
|
|
| 89 |
|
| 90 |
```python
|
| 91 |
from huggingface_hub import snapshot_download
|
| 92 |
+
import sys
|
| 93 |
|
| 94 |
d = snapshot_download("DylanCouzon/constella-zero")
|
| 95 |
sys.path.insert(0, d)
|
| 96 |
from zero_encoder import ZeroQueryEncoder
|
| 97 |
|
| 98 |
+
enc = ZeroQueryEncoder(d, variant="int8")
|
| 99 |
+
q_np = enc.encode(["how do mrna vaccines work?"])
|
| 100 |
+
assert np.abs(q_np[0] - q).max() < 1e-5
|
| 101 |
```
|
| 102 |
|
| 103 |
+
Zero and Nano share a document space, not a retrieval-quality guarantee: they have different
|
| 104 |
+
measured behavior, and interchangeability does not mean parity, equivalence, or a tie.
|
| 105 |
|
| 106 |
+
## How it works
|
|
|
|
|
|
|
|
|
|
| 107 |
|
| 108 |
+
Tokenize with WordPiece, special tokens on, no prefix, and truncation at 512 tokens. A token that
|
| 109 |
+
appears `c` times carries total weight `sqrt(c)`, so repetition saturates. The graph sums the rows,
|
| 110 |
+
divides by the weight sum, and L2-normalizes; an empty or near-zero bag falls back to normalized
|
| 111 |
+
`[CLS]`. Learned token weights are folded into the rows. This is still a bag of tokens: word order,
|
| 112 |
+
negation, and syntax are not represented.
|
| 113 |
|
| 114 |
+
The reported model is int8, which was within 0.00013 nDCG@10 of fp16.
|
|
|
|
| 115 |
|
| 116 |
## Files
|
| 117 |
|
| 118 |
+
| file | purpose | size |
|
| 119 |
+
|---|---|---:|
|
| 120 |
+
| `model.onnx` | FastEmbed/ONNX Runtime; pooled normalized `(batch, 1024)` output | 31 MB |
|
| 121 |
+
| `model_tokens.onnx` | token-level `(batch, sequence, 1024)` output | 31 MB |
|
| 122 |
+
| `model.npz` | NumPy reference path | 94 MB |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 123 |
|
| 124 |
+
Both graphs use opset 17 and standard operators. The int8 table is dequantized in-graph with one
|
| 125 |
+
fp32 scale per row. Tokenizer metadata enforces the frozen 512-token rule and dynamic padding.
|
|
|
|
|
|
|
| 126 |
|
| 127 |
## Results
|
| 128 |
|
| 129 |
+
The headline partition is **clean-4**: NFCorpus, SCIDOCS, SciFact, and TREC-COVID. ArguAna and
|
| 130 |
+
FiQA remain beside it with `†` because Stella discloses training/evaluation contact with them.
|
| 131 |
+
All values are exact-search nDCG@10.
|
| 132 |
|
| 133 |
+
| system | NFCorpus **(clean-4)** | SCIDOCS **(clean-4)** | SciFact **(clean-4)** | TREC-COVID **(clean-4)** | ArguAna† | FiQA† |
|
| 134 |
+
|---|---:|---:|---:|---:|---:|---:|
|
| 135 |
+
| constella-zero (int8) | 0.3124 | 0.1677 | 0.6101 | 0.5490 | 0.5916 | 0.3728 |
|
| 136 |
+
| constella-nano | 0.363080 | 0.217710 | 0.721097 | 0.787116 | 0.623296 | 0.477765 |
|
| 137 |
+
| BM25 | 0.3180 | 0.1565 | 0.6791 | 0.6099 | 0.4878 | 0.2532 |
|
| 138 |
+
| Stella teacher, symmetric | 0.4134 | 0.2395 | 0.7796 | 0.8234 | 0.6369 | 0.5536 |
|
|
|
|
| 139 |
|
| 140 |
+
† Stella-disclosed training/evaluation contact; excluded from clean-4.
|
|
|
|
| 141 |
|
| 142 |
+
Zero did **not** confirmatorily beat BM25. M7 C2 was +0.0165 across all six with raw 95% interval
|
| 143 |
+
[+0.0017, +0.0311], but its sign-flip p=0.0149 failed the Holm threshold of 0.0083. On clean-4,
|
| 144 |
+
the descriptive contrast was -0.0311 [-0.0517, -0.0109], with Zero at 0.4098 versus BM25 at
|
| 145 |
+
0.4409. Superiority is **UNESTABLISHED**.
|
| 146 |
|
| 147 |
+
The deployable hybrid recommendation and registered operator of record are distinct:
|
|
|
|
|
|
|
|
|
|
| 148 |
|
| 149 |
+
| Zero + BM25 fusion | prefetch | all-six macro | clean-4 macro |
|
| 150 |
+
|---|---:|---:|---:|
|
| 151 |
+
| Qdrant DBSF | 100 | 0.4887 | 0.4912 |
|
| 152 |
+
| M7 convex0 (`w=0.8`) | 1000 | 0.4911 | 0.4866 |
|
| 153 |
|
| 154 |
+
Both rows use `bm25s` with Lucene defaults as the lexical side, not Qdrant's own BM25, which has a
|
| 155 |
+
fixed `avg_len` and its own tokenizer — DBSF normalises over returned scores, so a different
|
| 156 |
+
lexical implementation shifts its inputs. For the DBSF row, each query's own document is excluded
|
| 157 |
+
*before* the prefetch is truncated to 100 (`must_not` on the point id); without that filter a
|
| 158 |
+
plain `limit: 100` spends a slot on the self-match. That affects only ArguAna (1,298 of 1,406
|
| 159 |
+
queries) and FiQA (55), so the clean-4 figures are unchanged either way.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 160 |
|
| 161 |
+
Use Qdrant DBSF at prefetch 100 in deployments. M7's convex0 is the registered operator of record,
|
| 162 |
+
but Qdrant does not implement it. No confidence interval compared these observations, so neither
|
| 163 |
+
superiority nor equivalence is established. M7 C3 likewise did not establish fusion superiority
|
| 164 |
+
over OpenSearch: +0.0043, raw 95% interval [-0.0063, +0.0151], p=0.219.
|
| 165 |
+
|
| 166 |
+
Full per-dataset fusion rows, registered contrasts, and source traces are in the
|
| 167 |
+
[M21 benchmark ledger](https://github.com/Dylancouzon/asymmetric-dual-encoders/blob/main/m21/BENCHMARKS.md).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 168 |
|
| 169 |
## Limits
|
| 170 |
|
| 171 |
+
- Reserved-four and BEIR-18 evidence is pending and unspent; the six datasets do not establish
|
| 172 |
+
broad domain coverage.
|
| 173 |
+
- Stella contact with ArguAna and FiQA makes all-six results secondary to clean-4.
|
| 174 |
+
- Zero is an English-only bag of tokens and truncates beyond 512 wordpieces.
|
| 175 |
+
- The document side is not cheap: the 400M-parameter Stella tower still runs at indexing time.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 176 |
|
| 177 |
## Costs
|
| 178 |
|
| 179 |
+
These are synthetic query latencies, not workload estimates. The common three-model protocol used
|
| 180 |
+
three fresh processes per model, batch one, four CPU threads; hydration includes imports,
|
| 181 |
+
verification, and load but excludes interpreter startup; first inference is separate; warm timing
|
| 182 |
+
uses five warmups and twenty 20-word samples; the OS disk cache was not flushed.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 183 |
|
| 184 |
+
| model | hydration | first query | warm 20-word p50 | peak RSS | model assets |
|
| 185 |
+
|---|---:|---:|---:|---:|---:|
|
| 186 |
+
| constella-zero | 0.2618 s | 0.3529 ms | 0.1119 ms | 275.4 MiB | 90.1 MiB |
|
| 187 |
+
| bge-small | 0.6726 s | 8.2401 ms | 6.8400 ms | 291.0 MiB | 127.6 MiB |
|
| 188 |
+
| constella-nano | 0.6907 s | 7.6685 ms | 7.2511 ms | 280.9 MiB | 132.3 MiB |
|
| 189 |
|
| 190 |
+
This measures query encoders only, not retrieval, ANN, or end-to-end system latency. The asset
|
| 191 |
+
column follows the common protocol; `model.onnx` itself is the 31 MB query graph listed above.
|
|
|
|
| 192 |
|
| 193 |
## Training
|
| 194 |
|
| 195 |
+
The table was trained by L2 regression against Stella query embeddings over 340,850 pairs plus
|
| 196 |
+
220,632 query-only rows from Amazon ESCI, FEVER, HotpotQA, SQuAD, NQ Open, TriviaQA, and Mr. TyDi
|
| 197 |
+
(English); MS MARCO was excluded. Wikipedia-derived sources retain CC BY-SA attribution; ESCI and
|
| 198 |
+
TriviaQA are Apache-2.0.
|
|
|
|
|
|
|
| 199 |
|
| 200 |
## Provenance
|
| 201 |
|
| 202 |
+
```text
|
| 203 |
+
run_id p35w-2m-s2500
|
| 204 |
+
table a7007b1a6af120b976f093fd69ddcb5001996ec0b84b5864b4fd25d7af878abf
|
| 205 |
+
teacher NovaSearch/stella_en_400M_v5 @ ffeb2b7ee715c226d4ffe5e4619f7dbb48624c20
|
| 206 |
+
comparator BAAI/bge-small-en-v1.5 @ 5c38ec7c405ec4b44b94cc5a9bb96e735b38267a
|
| 207 |
+
preproc prefix="" · special tokens · max_length=512 · pool_mode=sqrt
|
| 208 |
+
fingerprint adb24fb2e8cad66f
|
| 209 |
```
|
| 210 |
|
| 211 |
+
The weights are MIT licensed. The pinned Stella teacher/document tower and bge-small comparator
|
| 212 |
+
are also MIT and are attributed above.
|