Hypencoder β€” ONNX exports (fp32)

Ready-to-run fp32 ONNX artifacts for all four released Hypencoder checkpoints, so you can run them without PyTorch, transformers, or the paper's research package.

A Hypencoder replaces the fixed inner product of a bi-encoder with a learned, query-specific neural network: a hypernetwork turns the query into a small MLP (a "q-net") whose input is a document vector and whose output is the relevance score.

Exported from the upstream checkpoints (all Apache-2.0) with model2onnx.py from the Vespa sample app.

Why this exists. The upstream checkpoints cannot be loaded by transformers>=5: Hypencoder.__init__ calls AutoModel.from_pretrained() inside the constructor, and transformers 5 wraps cls(config) in a meta-device context, so the nested load is rejected by check_and_set_device_map. Exporting once and shipping ONNX removes torch, transformers, and the editable hypencoder-paper install requirement for running our sample app.

Contents

directory q-net blocks query-encoder outputs generated params per query passage_encoder query_encoder
2_layer/ 2 5 1,181,952 435.8 MB 480.8 MB
4_layer/ 4 9 2,363,136 435.8 MB 518.6 MB
6_layer/ 6 13 3,544,320 435.8 MB 556.4 MB
8_layer/ 8 17 4,725,504 435.8 MB 594.3 MB

Each directory holds passage_encoder.onnx, query_encoder.onnx and tokenizer.json.

The four passage encoders are not interchangeable β€” each checkpoint is a separate fine-tune, so the document towers differ. Pair a passage encoder only with the query encoder from the same directory.

I/O contract

Identical across depths except for the number of query-encoder outputs.

passage_encoder.onnx β€” input_ids, attention_mask (int64, [batch, seq]) β†’ last_hidden_state [batch, seq, 768]. CLS-pool it (take index 0) and do not L2-normalise.

query_encoder.onnx β€” input_ids, attention_mask (float32, [batch, seq]) β†’ W0 [768,768], b0 [768], … , W{n-1}, b{n-1}, Wout [768].

Inputs are float32 rather than int64 because Vespa tensors have no int64 cell type; the export wrapper casts internally. Weight matrices are pre-transposed to (out, in) so they arrive in Vespa's alphabetical dimension order.

Scoring

Each q-net block is linear β†’ ReLU β†’ parameter-free LayerNorm, with a residual connection, followed by a final linear projection to a scalar. For the 2-block model, as written in the Vespa sample app's rank profile:

import numpy as np

def layer_norm(v, eps=1e-5):
    return (v - v.mean(-1, keepdims=True)) / np.sqrt(v.var(-1, keepdims=True) + eps)

def score_2layer(x0, W0, b0, W1, b1, Wout):
    relu0 = np.maximum(x0 @ W0.T + b0, 0.0)
    res0  = layer_norm(relu0) + x0          # residual on the first block
    relu1 = np.maximum(res0 @ W1.T + b1, 0.0)
    return layer_norm(relu1) @ Wout         # no residual before the final projection

For the deeper checkpoints, treat q_net.py as authoritative for where residuals and layer norms are applied β€” the snippet above is written for the 2-block case only.

Usage with onnxruntime

import numpy as np, onnxruntime as ort
from tokenizers import Tokenizer

D = "2_layer"
tok  = Tokenizer.from_file(f"{D}/tokenizer.json")
enc  = ort.InferenceSession(f"{D}/passage_encoder.onnx")
qenc = ort.InferenceSession(f"{D}/query_encoder.onnx")

d = tok.encode("Mount Everest is Earth's highest mountain, at 8,849 metres.")
doc_vec = enc.run(None, {"input_ids": np.array([d.ids], dtype=np.int64),
                         "attention_mask": np.array([d.attention_mask], dtype=np.int64)}
                  )[0][:, 0]                                   # CLS pooling

q = tok.encode("tallest mountain in the world")
names = [o.name for o in qenc.get_outputs()]
w = dict(zip(names, qenc.run(None, {
    "input_ids": np.array([q.ids], dtype=np.float32),           # float, not int64
    "attention_mask": np.array([q.attention_mask], dtype=np.float32)})))

print(score_2layer(doc_vec.astype(np.float64),
                   w["W0"], w["b0"], w["W1"], w["b1"], w["Wout"]))

Usage with Vespa

The passage encoder and tokenizer can be fetched by URL at deploy time, so they never enter your application package:

<component id="passage_embedder" type="hugging-face-embedder">
  <transformer-model url="https://huggingface.co/andreer/hypencoder-onnx/resolve/main/2_layer/passage_encoder.onnx"/>
  <tokenizer-model  url="https://huggingface.co/andreer/hypencoder-onnx/resolve/main/2_layer/tokenizer.json"/>
  <pooling-strategy>cls</pooling-strategy>
  <normalize>false</normalize>
</component>

The query encoder must be a local file β€” onnx-model does not accept a URI (OnnxModel.setUri() throws "URI for ONNX models are not currently supported"), so download it into the package first:

mkdir -p app/models
curl -L -o app/models/query_encoder.onnx \
  https://huggingface.co/andreer/hypencoder-onnx/resolve/main/2_layer/query_encoder.onnx
onnx-model query_encoder {
    file: models/query_encoder.onnx
    input "input_ids": query(input_ids)
    input "attention_mask": query(attention_mask)
}

The sample app's rank profile implements 2 blocks only. Using 4_layer, 6_layer or 8_layer means extending that expression with the extra W{i}/b{i} blocks β€” the ONNX outputs are there, but the ranking expression is not written for them.

Modifications relative to the base models

Per Apache-2.0 Β§4(b): no weights were retrained or altered numerically. The export wrapper (a) returns the q-net weight tensors directly instead of a callable, (b) casts float input_ids/attention_mask to int64 internally, (c) pre-transposes weight matrices to (out, in), and (d) uses static layer-norm shapes so the legacy TorchScript exporter inlines weights into a single file rather than an external .data sidecar.

Citation

@inproceedings{killingback2025hypencoder,
  title     = {Hypencoder: Hypernetworks for Information Retrieval},
  author    = {Killingback, Julian and Zeng, Hansi and Zamani, Hamed},
  booktitle = {SIGIR},
  year      = {2025}
}

Weights Β© the original authors, Apache-2.0. This repository redistributes them in ONNX form under the same licence.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for andreer/hypencoder-onnx

Quantized
(1)
this model

Paper for andreer/hypencoder-onnx