Instructions to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("webmp3/Sakura-EmbeddingGemma-2-AutoRound") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Transformers
How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="webmp3/Sakura-EmbeddingGemma-2-AutoRound")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("webmp3/Sakura-EmbeddingGemma-2-AutoRound") model = AutoModel.from_pretrained("webmp3/Sakura-EmbeddingGemma-2-AutoRound", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Sakura EmbeddingGemma 2 — AutoRound W4A16
- Sakura EmbeddingGemma 2 — AutoRound W4A16
Sakura EmbeddingGemma 2 — AutoRound W4A16
Community Quantization. Not official from Google.
Quantized and evaluated by Sakura (webmp3) using Intel AutoRound optimization on top of Google's officialembeddinggemma-2architecture.
Overview
This repository provides an optimized AutoRound W4A16 (4-bit weights, 16-bit activations) quantization of google/embeddinggemma-2.
It is packaged as a standard Hugging Face repository containing packed Safetensors weights that run natively and out of the box with sentence-transformers and transformers.
- Upstream Model:
google/embeddinggemma-2 - Pinned Upstream Commit:
914f7f89142e33e77833254d9c9b90c3cef7303b - Base Architecture:
EmbeddingGemma2ForSequenceClassification/EmbeddingGemma2TextModel - Model File Size:
model.safetensors: 1,299,010,632 bytes / 1,299.01 MB / 1,238.83 MiB. The pinned BF16 original model file is 1,488,915,288 bytes / 1,488.92 MB / 1,419.94 MiB: 12.75% smaller on disk. This compares model files, not total directories or measured runtime memory. - Verification Status: Public Hub access verified; local release smoke test passed; packaged for Transformers & SentenceTransformers.
- Quantization Framework: AutoRound 0.16.0, as recorded in the published
quantization_config.json. - Quantization Scheme: W4A16 symmetric (
bits=4,group_size=64,iters=200), calibrated on 234 mixed retrieval texts (build/calibration_dual256.json); recorded packing formatauto_round:auto_gptq. - Version: v2 (2026-10-07). Re-quantized with a broader calibration set, more iterations and a smaller group size. The v1 release (group size 128, 100 iterations, 64 calibration texts) is kept in this repository's git history. v2 improves BF16 fidelity and external retrieval modestly; see the tables below.
- Release Decision: RELEASE GO (Early community release based on empirical fidelity retention)
Architectural Breakdown: What is W4 vs. What Remains BF16
We explicitly do not claim "Full INT4/Q4". High-fidelity embedding models require careful treatment of sensitive components:
| Component | Precision | Details |
|---|---|---|
| Text Backbone Linear Layers | W4A16 | 216 Linear layers in language_model.layers.0 through language_model.layers.23 (q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj). Group size 64, symmetric. |
| Token Embeddings | BF16 | language_model.embed_tokens (256,000 vocab) preserved in bfloat16 to avoid semantic vocabulary collapse. |
| Embedding Projection Head | BF16 | language_model.embedding_projection (768d output head) preserved in bfloat16 to preserve precise directional geometry. |
| Normalization Layers | BF16 | All RMSNorm and LayerNorm modules preserved in bfloat16 to prevent activation scale clipping. |
| Vision & Audio Towers | BF16 | Multimodal encoders (vision_tower, audio_tower) and multimodal projection heads are preserved unquantized. |
Fidelity & Benchmark Results
All evaluations were conducted against an unquantized bfloat16 reference across 100 query/document retrieval pairs (English and German) and 50 code retrieval pairs.
Release Gate Status
The candidate was evaluated against strict quality criteria:
- Empirical Gate Result: RELEASE GO / Early Community Release
- Gate Context: The initial theoretical target of $\ge 0.99$ Mean Cosine Similarity and $\ge 95%$ Top-5 Retrieval Agreement is not fully reached (v2 achieved: 0.9897 Mean Cosine, 89.0 % Top-5 Agreement; v1: 0.9878 / 87.6 %). With 100.0 % Top-1 Agreement, 100.0 % Recall@5 and 0.9814 Spearman Correlation, these benchmark results show retained retrieval success alongside measurable BF16-fidelity loss; no NaN or Inf anomaly was observed.
The companion GGUF HQ releases achieve higher reported BF16 fidelity on the same existing BF16 reference and text/code benchmark. They are separate quantization runs; their results do not change this Safetensors model's release-gate outcome.
| GGUF companion tier | MiB | Combined Mean Cosine (768d) | Combined Spearman (768d) | Text Top-5 | Text/Code Recall@5 |
|---|---|---|---|---|---|
| Q4 HQ | 169.79 | 0.99400270 | 0.98857239 | 90.00% | 100% / 100% |
| Q5 HQ — recommended balance | 199.88 | 0.99828082 | 0.99664205 | 95.40% | 100% / 100% |
| Q6 HQ | 231.75 | 0.99940187 | 0.99879852 | 96.80% | 100% / 100% |
Values are quoted from the GGUF README's Side-by-Side Comparison. Aggregation differs: this Safetensors section reports alignment over 200 text query/document vectors; the GGUF combined Mean Cosine/Spearman include all 300 text-and-code vectors. Text Top-5 uses the 100-pair text retrieval benchmark in both. These are not identically aggregated mean scores, and no all-metric comparison with Unsloth is implied.
MRL (Matryoshka Representation Learning) Dimensions
EmbeddingGemma 2 supports dimension truncation followed by L2-renormalization. The W4A16 model demonstrates robust stability across all standard MRL truncations:
| MRL Dimension | Mean Cosine Sim | Min Cosine Sim | Spearman Rank Corr | Mean L2 Drift | Top-1 Agreement | Top-5 Agreement | Recall@5 | nDCG@10 |
|---|---|---|---|---|---|---|---|---|
| 768d (Full) | 0.9897 | 0.9662 | 0.9814 | 0.1404 | 100.0 % | 89.0 % | 100.0 % | 0.9645 |
| 512d | 0.9899 | 0.9670 | 0.9815 | 0.1390 | 99.0 % | 88.6 % | 100.0 % | 0.9590 |
| 256d | 0.9910 | 0.9701 | 0.9832 | 0.1309 | 100.0 % | 85.6 % | 100.0 % | 0.9579 |
| 128d | 0.9934 | 0.9807 | 0.9828 | 0.1120 | 99.0 % | 86.6 % | 100.0 % | 0.9454 |
Language & Domain Subsets (768d)
| Subset | Mean Cosine Sim | Top-1 Agreement | Top-5 Agreement | Recall@5 |
|---|---|---|---|---|
| English Queries & Docs | 0.9897 | 100.0 % | 88.4 % | 100.0 % |
| German Queries & Docs | 0.9898 | 100.0 % | 87.6 % | 100.0 % |
| Code Retrieval | 0.9893 | 100.0 % | 83.2 % | 100.0 % |
External retrieval (BEIR / MTEB test sets, measured)
Same texts and prefixes for all rows, texts truncated to 128 tokens; 500-document corpora (all positives + sampled distractors). BF16 = unquantized google/embeddinggemma-2.
| Benchmark | Model | MRR | nDCG@10 | Recall@10 | MRR vs BF16 | nDCG@10 vs BF16 |
|---|---|---|---|---|---|---|
| SciFact (300 queries) | BF16 reference | 0.8971 | 0.9137 | 0.9700 | 100 % | 100 % |
| W4A16 v2 (this release) | 0.8921 | 0.9105 | 0.9733 | 99.4 % | 99.7 % | |
| W4A16 v1 (previous release) | 0.8892 | 0.9051 | 0.9667 | 99.1 % | 99.1 % | |
| NFCorpus (249 queries) | BF16 reference | 0.4609 | 0.3099 | 0.6386 | 100 % | 100 % |
| W4A16 v2 (this release) | 0.4509 | 0.2971 | 0.6265 | 97.8 % | 95.9 % | |
| W4A16 v1 (previous release) | 0.4466 | 0.2973 | 0.6145 | 96.9 % | 95.9 % |
The differences between v1 and v2 are small and not uniform. v2 is better in mean cosine and Spearman correlation in every MRL dimension and subset, in Top-5 agreement at 768d (+1.4 points), 128d, English and German, and in SciFact/NFCorpus MRR. It is slightly worse in Top-5 agreement at 256d (-0.6 points) and on the code subset (-4.0 points: 83.2 % vs. 87.2 %), and nDCG@10 on the text benchmark is 0.001-0.003 lower; NFCorpus nDCG@10 is unchanged. NFCorpus has only 249 queries, so differences of about one point are within noise. Overall v2 is a modest improvement, not a step change.
Runtime & Community Format Comparison
| Distribution / Format | Runtime Compatibility | Direct Transformers / ST Support | Notes |
|---|---|---|---|
| Sakura AutoRound W4A16 (This repo) | Python, PyTorch, Transformers, SentenceTransformers | Yes (Plug-and-play) | Runs directly in existing Python AI pipelines without needing custom binary builds. |
GGUF Community Releases (unsloth, ggml-org) |
llama.cpp |
No (requires llama-server or bindings) | EmbeddingGemma 2 support is upstream in llama.cpp via PR #30054, merged 2026-10-06. The earlier unknown model architecture result came from an older local build; use a build that includes this support. Community GGUF controls have since been loaded and benchmarked with a supported build. |
ONNX Community Releases (onnx-community) |
ONNX Runtime / Transformers.js | No (ONNX graph format) | Modular multi-graph export tailored for WebGPU/JavaScript execution. |
| Sakura AutoRound GGUF HQ (Companion) | llama.cpp |
Via llama-server / bindings | Q4 HQ 169.79 MiB, Q5 HQ 199.88 MiB (recommended balance), Q6 HQ 231.75 MiB. This Safetensors/ST package is for native Python pipelines; the standalone text GGUF variants are for llama.cpp. |
Safetensors vs. GGUF HQ: separate quantization runs
These are separate optimization runs, not two exports of one quantized state. The published Safetensors quantization_config.json records a different scheme and iteration count from the preserved GGUF cal256 build configurations/logs. Both use AutoRound 0.16.0 and the same base model, but that does not imply identical optimized weights or calibration.
| Setting | This Safetensors W4A16 release | GGUF HQ cal256 runs |
|---|---|---|
| Transformer weight types | 4-bit symmetric W4A16 | Q4_K/Q5_K/Q6_K base with 48 Q8_0 block-PLE linears |
| Weight grouping | Group size 64 | Q4_K/Q5_K: 32 weights per subgroup, 8 subgroups per 256-weight superblock; Q6_K: 16, 16 subgroups per 256-weight superblock. Q8_0 uses 32-weight blocks. |
| Iterations | 200, published configuration | 50, saved build configurations/logs |
| Calibration selection | 234 real retrieval texts from 13 public datasets (queries with the EmbeddingGemma query prefix, passages with the document prefix; EN/DE/code/science/web), manifest build/calibration_dual256.json (SHA256 d245f6268fe39f9fca92befa80ee709f5108852a880dc7a6714bc6e772de0ecd), built with build/requant.py |
256 synthetic retrieval samples: 48 EN queries, 48 EN docs, 48 DE queries, 48 DE docs, 32 code queries and 32 code snippets; zero exact benchmark overlap |
| Optimization/export path | AutoRound W4A16; published packing format auto_round:auto_gptq; Safetensors for Transformers/ST |
Native AutoRound SignRoundV2 (enable_alg_ext=True), matching GGUF optimized-state packing, then selected embedding-path precision overrides |
| AutoRound version | 0.16.0, published configuration | 0.16.0, preserved builds |
| Scope | Full multimodal checkpoint, vision/audio and sensitive components retained in BF16 | Standalone text GGUF; CPU/native llama.cpp runtime |
Provenance of this release (v2): the calibration texts, the quantization script (build/requant.py, run with AutoRound 0.16.0 on CPU) and the evaluation results are published or reproducible from this repository. The v1 release (git history) was built by the earlier pipeline script with up to 64 texts from a local calibration file whose exact historical content was not preserved. The calibration texts come from public datasets (some, e.g. MS MARCO, carry research-use terms); only the texts used for calibration are included.
The GGUF cal256 runs do retain explicit sample count, calibration SHA256 b65d22bad14722ab02d316ddae2ea6f860b92c94669f5717321b66e81abb91b8, resolved per-layer configuration, native optimized-state packing evidence and training logs. These differences are enough to rule out a shared optimized quantization state. Public GGUF settings and runtime provenance are linked from the companion README; the inspected W4A16 pipeline settings are summarized here without claiming historical inputs that were not preserved.
Quickstart & Usage
1. With SentenceTransformers (Recommended)
from sentence_transformers import SentenceTransformer
import torch
# Load the quantized model
model = SentenceTransformer(
"webmp3/Sakura-EmbeddingGemma-2-AutoRound",
model_kwargs={"torch_dtype": torch.bfloat16}
)
# Text Retrieval Query (using official prompt_name)
query = "What is quantum entanglement?"
query_embedding = model.encode(query, prompt_name="SearchQuery")
# Documents (unprompted)
docs = [
"Quantum entanglement is a phenomenon where particles remain connected regardless of distance.",
"The recipe for chocolate chip cookies requires flour, butter, and sugar."
]
doc_embeddings = model.encode(docs)
# Compute similarity
similarities = model.similarity(query_embedding, doc_embeddings)
print("Similarities:", similarities)
2. Matryoshka Dimension Truncation (MRL)
To reduce memory and storage footprint, simply slice the vector and re-normalize:
import torch
import torch.nn.functional as F
# 128-dimensional embedding
full_embedding = model.encode(["Example sentence"], convert_to_tensor=True)
mrl_128 = full_embedding[:, :128]
mrl_128_normalized = F.normalize(mrl_128, p=2, dim=-1)
print("128d shape:", mrl_128_normalized.shape)
3. With Hugging Face Transformers
from transformers import AutoModel, AutoTokenizer
import torch
model_id = "webmp3/Sakura-EmbeddingGemma-2-AutoRound"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, torch_dtype=torch.bfloat16)
inputs = tokenizer(["SearchQuery: What is machine learning?"], return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
Limitations & Honest Disclosure
- Text Precision and File Size: The 216 text-backbone linear layers use 4-bit packed weights.
model.safetensorsis 1,299,010,632 bytes (1,299.01 MB / 1,238.83 MiB), compared with 1,488,915,288 bytes (1,488.92 MB / 1,419.94 MiB) for the pinned full BF16 original. The 12.75% model-file saving is modest: token embeddings and multimodal towers remain BF16. These on-disk measurements do not establish a runtime-memory reduction; the former ~5.5 GB BF16 comparison is not supported by the checked model files. - Multimodal Towers: The vision and audio towers remain in 16-bit precision. If you do not use vision or audio inputs, memory consumption can be minimized by only loading the text language model.
- Execution Device: Optimal execution requires modern CPU (AVX-512 / VNNI) or GPU environments supporting accelerated bfloat16 and packed int4 kernels.
SHA256 Checksums
Every file in this release has been cryptographically verified:
1be8f9083684a77b24317163310782efd0df30e5ce5cc2c097d447b3a5402ce1 model.safetensors
48d4e0ca37c329ae9fc6582f7de5224f6d5ce9a3b81ea00dfe61eaf7a622c45c config.json
072b3e5dc502e1beabac6414e8a663516995b75abe1c03fedf1ddb696eb49483 quantization_config.json
031e56a498d33c349ab489a21885bcfe25b4fcba841149dc99e1e90d4a7c28f5 config_sentence_transformers.json
b1bcd9f2dce3ae863b359e87d0710b5dbc3314a59ecb4e2f97c7778fc8e4b228 sentence_bert_config.json
3d02572a0455b832de67fb8e63a54981bc7e8b46e337c95e917bd8122a533bfd modules.json
ea2ae257e901064abdd98dceb19f2b0da06af600bed15e0f99f5c85c37ee9d78 preprocessor_config.json
168f6a08522f3ce5dea596d94d003af2fd691742d4f41fe1f9d8cce76bfbf69c processor_config.json
4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4 tokenizer.json
17bd5d6e9364ca49a534e1502076593317c298d4a663623091ed45388f004874 tokenizer_config.json
4b852efc0b9960283e735363331e6f325b33bc74bdbaa076f595bc4e9b94d85e chat_template.jinja
8759bdf7c77efc7df7723f64856a593c8943b71ee38baf2a88771fbaf78438f9 1_Pooling/config.json
cdb09dfca347a56aa2d691744e38d5ad3c7cbc2834e7181272b9a15328b82524 2_Normalize/config.json
d245f6268fe39f9fca92befa80ee709f5108852a880dc7a6714bc6e772de0ecd build/calibration_dual256.json
Citation & Acknowledgements
- Upstream Model by Google: google/embeddinggemma-2
- AutoRound framework by Intel: Intel AutoRound
- Quantization & Benchmark Pipeline by Sakura (
webmp3).
中文说明 · 樱花 (Simplified Chinese)
English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。
Sakura EmbeddingGemma 2 — AutoRound W4A16
社区量化版本。并非 Google 官方发布。
由 Sakura(webmp3)在 Google 官方embeddinggemma-2架构之上,使用 Intel AutoRound 优化进行量化和评估。
概览
本仓库提供 google/embeddinggemma-2 的优化 AutoRound W4A16(4 比特权重,16 比特激活)量化版本。
它被打包为标准的 Hugging Face 仓库,包含打包好的 Safetensors 权重,可以在 sentence-transformers 和 transformers 中 原生地、开箱即用地 运行。
- 上游模型:
google/embeddinggemma-2 - 固定的上游提交:
914f7f89142e33e77833254d9c9b90c3cef7303b - 基础架构:
EmbeddingGemma2ForSequenceClassification/EmbeddingGemma2TextModel - 模型文件大小:
model.safetensors:1,299,010,632 字节 / 1,299.01 MB / 1,238.83 MiB。固定的 BF16 原始模型文件为 1,488,915,288 字节 / 1,488.92 MB / 1,419.94 MiB:**磁盘上小 12.75%**。这比较的是模型文件,而不是整个目录或实测的运行时内存。 - 验证状态: 已验证可公开访问 Hub;本地发布冒烟测试通过;已为 Transformers 和 SentenceTransformers 打包。
- 量化框架: AutoRound 0.16.0,记录在已发布的
quantization_config.json中。 - 量化方案: W4A16 对称(
bits=4、group_size=64、iters=200),在 234 条混合检索文本(build/calibration_dual256.json)上校准;记录的打包格式为auto_round:auto_gptq。 - 版本: **v2(2026-10-07)**。使用更广泛的校准集、更多的迭代次数和更小的分组大小重新量化。v1 版本(分组大小 128,迭代 100 次,64 条校准文本)保留在本仓库的 git 历史中。v2 对 BF16 保真度和外部检索有适度改善;见下表。
- 发布决定: RELEASE GO(基于实测保真度保持情况的早期社区发布)
架构拆解:哪些是 W4,哪些保持 BF16
我们明确 不 声称是“完整的 INT4/Q4”。高保真的嵌入模型需要谨慎对待敏感组件:
| 组件 | 精度 | 详情 |
|---|---|---|
| 文本主干线性层 | W4A16 | language_model.layers.0 到 language_model.layers.23 中的 216 个线性层(q_proj、k_proj、v_proj、o_proj、gate_proj、up_proj、down_proj)。分组大小 64,对称。 |
| Token 嵌入 | BF16 | language_model.embed_tokens(256,000 词表)保持 bfloat16,以避免语义词表坍缩。 |
| 嵌入投影头 | BF16 | language_model.embedding_projection(768 维输出头)保持 bfloat16,以保留精确的方向几何。 |
| 归一化层 | BF16 | 所有 RMSNorm 和 LayerNorm 模块保持 bfloat16,以防止激活尺度被截断。 |
| 视觉与音频塔 | BF16 | 多模态编码器(vision_tower、audio_tower)和多模态投影头保持未量化。 |
保真度与基准结果
所有评估都是针对未量化的 bfloat16 参照 进行的,涵盖 100 个查询/文档检索对(英语和德语)以及 50 个代码检索对。
发布关卡状态
候选版本依据严格的质量标准进行了评估:
- 实测关卡结果: RELEASE GO / 早期社区发布
- 关卡背景: 最初的理论目标是 $\ge 0.99$ 的 Mean Cosine Similarity 和 $\ge 95%$ 的 Top-5 检索一致率,但并未完全达到(v2 达到:0.9897 Mean Cosine,89.0 % Top-5 一致率;v1:0.9878 / 87.6 %)。100.0 % 的 Top-1 一致率、100.0 % 的 Recall@5 和 0.9814 的 Spearman 相关系数 表明,这些基准结果显示检索成功率得以保持,同时存在可测量的 BF16 保真度损失;没有观察到 NaN 或 Inf 异常。
配套的 GGUF HQ 发布版本 在相同的现有 BF16 参照和文本/代码基准上取得了更高的 BF16 保真度。它们是独立的量化运行;其结果 不会 改变这个 Safetensors 模型的发布关卡结论。
| GGUF 配套档位 | MiB | 综合 Mean Cosine (768d) | 综合 Spearman (768d) | Text Top-5 | Text/Code Recall@5 |
|---|---|---|---|---|---|
| Q4 HQ | 169.79 | 0.99400270 | 0.98857239 | 90.00% | 100% / 100% |
| Q5 HQ — 推荐的平衡之选 | 199.88 | 0.99828082 | 0.99664205 | 95.40% | 100% / 100% |
| Q6 HQ | 231.75 | 0.99940187 | 0.99879852 | 96.80% | 100% / 100% |
数值引自 GGUF README 的 并排比较。汇总方式不同: 本 Safetensors 部分报告的是 200 个文本查询/文档向量上的对齐程度;GGUF 的综合 Mean Cosine/Spearman 包含全部 300 个文本和代码向量。两者的 Text Top-5 都使用 100 对的文本检索基准。这些不是以相同方式汇总的平均分,也不暗示与 Unsloth 在所有指标上的比较。
MRL(Matryoshka 表示学习)维度
EmbeddingGemma 2 支持先截断维度、再做 L2 重新归一化。W4A16 模型在所有标准 MRL 截断下都表现出稳健的稳定性:
| MRL 维度 | Mean Cosine Sim | Min Cosine Sim | Spearman Rank Corr | Mean L2 Drift | Top-1 一致率 | Top-5 一致率 | Recall@5 | nDCG@10 |
|---|---|---|---|---|---|---|---|---|
| 768d (完整) | 0.9897 | 0.9662 | 0.9814 | 0.1404 | 100.0 % | 89.0 % | 100.0 % | 0.9645 |
| 512d | 0.9899 | 0.9670 | 0.9815 | 0.1390 | 99.0 % | 88.6 % | 100.0 % | 0.9590 |
| 256d | 0.9910 | 0.9701 | 0.9832 | 0.1309 | 100.0 % | 85.6 % | 100.0 % | 0.9579 |
| 128d | 0.9934 | 0.9807 | 0.9828 | 0.1120 | 99.0 % | 86.6 % | 100.0 % | 0.9454 |
语言与领域子集 (768d)
| 子集 | Mean Cosine Sim | Top-1 一致率 | Top-5 一致率 | Recall@5 |
|---|---|---|---|---|
| 英文查询与文档 | 0.9897 | 100.0 % | 88.4 % | 100.0 % |
| 德文查询与文档 | 0.9898 | 100.0 % | 87.6 % | 100.0 % |
| 代码检索 | 0.9893 | 100.0 % | 83.2 % | 100.0 % |
外部检索(BEIR / MTEB 测试集,实测)
所有行使用相同的文本和前缀,文本截断为 128 个 token;500 篇文档的语料(所有正例 + 采样的干扰项)。BF16 = 未量化的 google/embeddinggemma-2。
| 基准 | 模型 | MRR | nDCG@10 | Recall@10 | 相对 BF16 的 MRR | 相对 BF16 的 nDCG@10 |
|---|---|---|---|---|---|---|
| SciFact(300 个查询) | BF16 参照 | 0.8971 | 0.9137 | 0.9700 | 100 % | 100 % |
| W4A16 v2(本次发布) | 0.8921 | 0.9105 | 0.9733 | 99.4 % | 99.7 % | |
| W4A16 v1(上一次发布) | 0.8892 | 0.9051 | 0.9667 | 99.1 % | 99.1 % | |
| NFCorpus(249 个查询) | BF16 参照 | 0.4609 | 0.3099 | 0.6386 | 100 % | 100 % |
| W4A16 v2(本次发布) | 0.4509 | 0.2971 | 0.6265 | 97.8 % | 95.9 % | |
| W4A16 v1(上一次发布) | 0.4466 | 0.2973 | 0.6145 | 96.9 % | 95.9 % |
v1 与 v2 之间的差异很小,而且并不一致。v2 在每个 MRL 维度和子集的 mean cosine 和 Spearman 相关性上更好,在 768d(+1.4 个百分点)、128d、英文和德文的 Top-5 一致率上更好,在 SciFact/NFCorpus 的 MRR 上也更好。它在 256d 的 Top-5 一致率(-0.6 个百分点)和代码子集(-4.0 个百分点:83.2 % 对 87.2 %)上略差,文本基准的 nDCG@10 低 0.001-0.003;NFCorpus 的 nDCG@10 没有变化。NFCorpus 只有 249 个查询,因此约一个百分点的差异处于噪声范围内。总体而言,v2 是适度的改进,不是一次飞跃。
运行时与社区格式对比
| 发行版 / 格式 | 运行时兼容性 | 是否直接支持 Transformers / ST | 说明 |
|---|---|---|---|
| Sakura AutoRound W4A16 (本仓库) | Python、PyTorch、Transformers、SentenceTransformers | 是(即插即用) | 可直接在现有的 Python AI 流水线中运行,无需自定义二进制构建。 |
GGUF 社区发布版本(unsloth、ggml-org) |
llama.cpp |
否(需要 llama-server 或绑定) | EmbeddingGemma 2 在 llama.cpp 中通过 PR #30054 获得上游支持,于 2026-10-06 合并。较早出现的 unknown model architecture 结果来自较旧的本地构建;请使用包含此支持的构建。此后,社区 GGUF 对照版本已用受支持的构建加载并做了基准测试。 |
ONNX 社区发布版本(onnx-community) |
ONNX Runtime / Transformers.js | 否(ONNX 图格式) | 为 WebGPU/JavaScript 执行量身定制的模块化多图导出。 |
| Sakura AutoRound GGUF HQ (配套) | llama.cpp |
通过 llama-server / 绑定 | Q4 HQ 169.79 MiB,Q5 HQ 199.88 MiB(推荐的平衡之选),Q6 HQ 231.75 MiB。本 Safetensors/ST 包适用于原生 Python 流水线;独立的文本 GGUF 变体适用于 llama.cpp。 |
Safetensors 与 GGUF HQ:独立的量化运行
这些是独立的优化运行,而不是同一个量化状态的两种导出。 已发布的 Safetensors quantization_config.json 所记录的方案和迭代次数,与保存下来的 GGUF cal256 构建配置/日志不同。两者都使用 AutoRound 0.16.0 和相同的基础模型,但这并不意味着优化后的权重或校准相同。
| 设置 | 本 Safetensors W4A16 发布版本 | GGUF HQ cal256 运行 |
|---|---|---|
| Transformer 权重类型 | 4 比特对称 W4A16 | Q4_K/Q5_K/Q6_K 基础,加 48 个 Q8_0 block-PLE 线性层 |
| 权重分组 | 分组大小 64 | Q4_K/Q5_K:每个子组 32 个权重,每个 256 权重的超级块含 8 个子组;Q6_K:16,每个 256 权重的超级块含 16 个子组。Q8_0 使用 32 权重的块。 |
| 迭代次数 | 200,已发布的配置 | 50,已保存的构建配置/日志 |
| 校准选择 | 来自 13 个公开数据集的 234 条真实检索文本(带 EmbeddingGemma 查询前缀的查询、带文档前缀的段落;EN/DE/代码/科学/网页),清单 build/calibration_dual256.json(SHA256 d245f6268fe39f9fca92befa80ee709f5108852a880dc7a6714bc6e772de0ecd),由 build/requant.py 构建 |
256 个合成检索样本:48 个英文查询、48 个英文文档、48 个德文查询、48 个德文文档、32 个代码查询和 32 个代码片段;与基准完全没有重合 |
| 优化/导出路径 | AutoRound W4A16;已发布的打包格式 auto_round:auto_gptq;供 Transformers/ST 使用的 Safetensors |
原生 AutoRound SignRoundV2(enable_alg_ext=True),匹配的 GGUF 优化状态打包,然后是选定的嵌入路径精度覆盖 |
| AutoRound 版本 | 0.16.0,已发布的配置 | 0.16.0,保存下来的构建 |
| 范围 | 完整的多模态检查点,视觉/音频和敏感组件保持 BF16 | 独立的文本 GGUF;CPU/原生 llama.cpp 运行时 |
本次发布(v2)的来源: 校准文本、量化脚本(build/requant.py,在 CPU 上使用 AutoRound 0.16.0 运行)和评估结果已发布,或可从本仓库复现。v1 版本(git 历史)由较早的流水线脚本构建,使用来自一个本地校准文件的最多 64 条文本,该文件的确切历史内容没有被保留。校准文本来自公开数据集(其中一些,例如 MS MARCO,带有仅限研究使用的条款);只包含用于校准的文本。
GGUF cal256 运行确实保留了明确的样本数量、校准 SHA256 b65d22bad14722ab02d316ddae2ea6f860b92c94669f5717321b66e81abb91b8、已解析的逐层配置、原生优化状态打包的证据以及训练日志。这些差异足以排除共享同一个优化后的量化状态。公开的 GGUF 设置和运行时来源已在配套 README 中链接;经过检查的 W4A16 流水线设置在此处做了概述,而没有声称那些未被保留的历史输入。
快速开始与用法
1. 使用 SentenceTransformers(推荐)
from sentence_transformers import SentenceTransformer
import torch
# Load the quantized model
model = SentenceTransformer(
"webmp3/Sakura-EmbeddingGemma-2-AutoRound",
model_kwargs={"torch_dtype": torch.bfloat16}
)
# Text Retrieval Query (using official prompt_name)
query = "What is quantum entanglement?"
query_embedding = model.encode(query, prompt_name="SearchQuery")
# Documents (unprompted)
docs = [
"Quantum entanglement is a phenomenon where particles remain connected regardless of distance.",
"The recipe for chocolate chip cookies requires flour, butter, and sugar."
]
doc_embeddings = model.encode(docs)
# Compute similarity
similarities = model.similarity(query_embedding, doc_embeddings)
print("Similarities:", similarities)
2. Matryoshka 维度截断(MRL)
要减小内存和存储占用,只需切片向量并重新归一化:
import torch
import torch.nn.functional as F
# 128-dimensional embedding
full_embedding = model.encode(["Example sentence"], convert_to_tensor=True)
mrl_128 = full_embedding[:, :128]
mrl_128_normalized = F.normalize(mrl_128, p=2, dim=-1)
print("128d shape:", mrl_128_normalized.shape)
3. 使用 Hugging Face Transformers
from transformers import AutoModel, AutoTokenizer
import torch
model_id = "webmp3/Sakura-EmbeddingGemma-2-AutoRound"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, torch_dtype=torch.bfloat16)
inputs = tokenizer(["SearchQuery: What is machine learning?"], return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
局限与坦率的披露
- 文本精度与文件大小: 216 个文本主干线性层使用 4 比特打包权重。
model.safetensors为 1,299,010,632 字节(1,299.01 MB / 1,238.83 MiB),而固定的完整 BF16 原始文件为 1,488,915,288 字节(1,488.92 MB / 1,419.94 MiB)。12.75% 的模型文件节省是有限的:token 嵌入和多模态塔仍为 BF16。这些磁盘上的测量并不能证明运行时内存有所减少;此前提到的约 5.5 GB 的 BF16 对比,在核对过的模型文件上并不成立。 - 多模态塔: 视觉和音频塔保持 16 比特精度。如果你不使用视觉或音频输入,只加载文本语言模型即可把内存消耗降到最低。
- 执行设备: 最佳执行需要现代 CPU(AVX-512 / VNNI)或支持加速 bfloat16 和打包 int4 内核的 GPU 环境。
SHA256 校验和
本次发布中的每个文件都已经过加密校验:
1be8f9083684a77b24317163310782efd0df30e5ce5cc2c097d447b3a5402ce1 model.safetensors
48d4e0ca37c329ae9fc6582f7de5224f6d5ce9a3b81ea00dfe61eaf7a622c45c config.json
072b3e5dc502e1beabac6414e8a663516995b75abe1c03fedf1ddb696eb49483 quantization_config.json
031e56a498d33c349ab489a21885bcfe25b4fcba841149dc99e1e90d4a7c28f5 config_sentence_transformers.json
b1bcd9f2dce3ae863b359e87d0710b5dbc3314a59ecb4e2f97c7778fc8e4b228 sentence_bert_config.json
3d02572a0455b832de67fb8e63a54981bc7e8b46e337c95e917bd8122a533bfd modules.json
ea2ae257e901064abdd98dceb19f2b0da06af600bed15e0f99f5c85c37ee9d78 preprocessor_config.json
168f6a08522f3ce5dea596d94d003af2fd691742d4f41fe1f9d8cce76bfbf69c processor_config.json
4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4 tokenizer.json
17bd5d6e9364ca49a534e1502076593317c298d4a663623091ed45388f004874 tokenizer_config.json
4b852efc0b9960283e735363331e6f325b33bc74bdbaa076f595bc4e9b94d85e chat_template.jinja
8759bdf7c77efc7df7723f64856a593c8943b71ee38baf2a88771fbaf78438f9 1_Pooling/config.json
cdb09dfca347a56aa2d691744e38d5ad3c7cbc2834e7181272b9a15328b82524 2_Normalize/config.json
d245f6268fe39f9fca92befa80ee709f5108852a880dc7a6714bc6e772de0ecd build/calibration_dual256.json
引用与致谢
- Google 的上游模型:google/embeddinggemma-2
- Intel 的 AutoRound 框架:Intel AutoRound
- 量化与基准流水线由 Sakura(
webmp3)完成。
- Downloads last month
- 35
Model tree for webmp3/Sakura-EmbeddingGemma-2-AutoRound
Base model
google/embeddinggemma-2