Geralt-Targaryen commited on
Commit
2f1beed
·
verified ·
1 Parent(s): b67bc89

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +23 -0
README.md CHANGED
@@ -115,6 +115,14 @@ F2LLM-v2 is fully open. We release base models in 5 sizes, instruct models in 8
115
  | 8B | [🤗F2LLM-v2-8B-Preview](https://huggingface.co/codefuse-ai/F2LLM-v2-8B-Preview) | [🤗F2LLM-v2-8B](https://huggingface.co/codefuse-ai/F2LLM-v2-8B) |
116
  | 14B | [🤗F2LLM-v2-14B-Preview](https://huggingface.co/codefuse-ai/F2LLM-v2-14B-Preview) | [🤗F2LLM-v2-14B](https://huggingface.co/codefuse-ai/F2LLM-v2-14B) |
117
 
 
 
 
 
 
 
 
 
118
  ## Usage
119
 
120
  ### With Sentence Transformers
@@ -199,6 +207,21 @@ In general, for retrieval and reranking tasks:
199
 
200
  For symmetric tasks such as STS, clustering, and bitext mining, you can encode the documents either with or without prompts. The model is trained to support both scenarios.
201
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
202
  ## Intermediate Checkpoints
203
 
204
  To facilitate future research, we release intermediate checkpoints in the `intermediate_checkpoints` branch.
 
115
  | 8B | [🤗F2LLM-v2-8B-Preview](https://huggingface.co/codefuse-ai/F2LLM-v2-8B-Preview) | [🤗F2LLM-v2-8B](https://huggingface.co/codefuse-ai/F2LLM-v2-8B) |
116
  | 14B | [🤗F2LLM-v2-14B-Preview](https://huggingface.co/codefuse-ai/F2LLM-v2-14B-Preview) | [🤗F2LLM-v2-14B](https://huggingface.co/codefuse-ai/F2LLM-v2-14B) |
117
 
118
+ ## Performance
119
+
120
+ The F2LLM-v2 family set a new state-of-the-art on a wide range of MTEB benchmarks, including Code, European, Scandinavian, German, French, Spanish, Polish, Dutch, Japanese, Vietnamese, Thai, Indic, Persian, among others.
121
+
122
+ <img src="img/performance.png" width="100%" alt="Performance">
123
+
124
+ For details, refer to the [MTEB leaderboard](https://huggingface.co/spaces/mteb/leaderboard).
125
+
126
  ## Usage
127
 
128
  ### With Sentence Transformers
 
207
 
208
  For symmetric tasks such as STS, clustering, and bitext mining, you can encode the documents either with or without prompts. The model is trained to support both scenarios.
209
 
210
+ ### MRL Support
211
+
212
+ This model is trained with Matryoshka Representation Learning (MRL), allowing for a superior tradeoff between performance and embedding size. You can truncate the embeddings to keep only the first `d` dimensions to reduce storage and speed up vector search in downstream systems. The model is trained with a smallest Matryoshka dimension of 8.
213
+
214
+ <img src="img/mrl.png" width="80%" alt="MRL results">
215
+
216
+ Example:
217
+ ```python
218
+ embedding = embedding[..., :128]
219
+ embedding = torch.nn.functional.normalize(embedding, p=2, dim=-1)
220
+ ```
221
+ > Note: you need to apply normalization **after** trucation, not the other way around.
222
+
223
+ If you are using sentence transformer, you can also simply pass `truncate_dim=128` to the encode interface.
224
+
225
  ## Intermediate Checkpoints
226
 
227
  To facilitate future research, we release intermediate checkpoints in the `intermediate_checkpoints` branch.