mlabonne v4zhong commited on
Commit
73c4489
·
1 Parent(s): aa6734b

GGUF release: F16/Q8_0/Q4_0, card + fill-mask.py (stock llama.cpp usage) (#2)

Browse files

- GGUF release: F16/Q8_0/Q4_0, card + fill-mask.py (stock llama.cpp usage) (e47fa578f845129f8915b6aad81c719d65c19510)


Co-authored-by: v4zhong <v4zhong@users.noreply.huggingface.co>

Files changed (2) hide show
  1. README.md +11 -4
  2. fill-mask.py +52 -0
README.md CHANGED
@@ -59,10 +59,18 @@ Find more information about LFM2.5-Encoder-230M in our [blog post](https://www.l
59
 
60
  Example usage with [llama.cpp](https://github.com/ggml-org/llama.cpp):
61
 
62
- Run masked-token prediction — the prompt's literal `[MASK]` is replaced with the model's mask token, one non-causal forward pass is run, and the top-K predictions at the mask position are printed:
 
 
 
 
 
 
63
 
64
  ```bash
65
- llama-fill-mask -hf LiquidAI/LFM2.5-Encoder-230M-GGUF "The capital of France is [MASK]." 5
 
 
66
  # 1 16.17 ' Paris'
67
  # 2 13.41 ' Strasbourg'
68
  # 3 13.35 'Paris'
@@ -70,10 +78,9 @@ llama-fill-mask -hf LiquidAI/LFM2.5-Encoder-230M-GGUF "The capital of France is
70
  # 5 11.87 ' Versailles'
71
  ```
72
 
73
- For downstream use, the encoder body also serves per-token embeddings:
74
 
75
  ```bash
76
- llama-server -hf LiquidAI/LFM2.5-Encoder-230M-GGUF --embeddings
77
  curl -s http://localhost:8080/embedding -d '{"content": "hello world"}'
78
  ```
79
 
 
59
 
60
  Example usage with [llama.cpp](https://github.com/ggml-org/llama.cpp):
61
 
62
+ Start llama-server with per-token embeddings
63
+ ```bash
64
+ hf download LiquidAI/LFM2.5-Encoder-230M-GGUF LFM2.5-Encoder-230M-F16.gguf --local-dir .
65
+ llama-server -m LFM2.5-Encoder-230M-F16.gguf --embeddings --pooling none
66
+ ```
67
+
68
+ Run masked-token prediction — the mask position's logits come from the per-token hidden states and the tied embedding matrix read from the GGUF ([`fill-mask.py`](./fill-mask.py) in this repo)
69
 
70
  ```bash
71
+ ❯ uv run fill-mask.py LFM2.5-Encoder-230M-F16.gguf "The capital of France is [MASK]."
72
+
73
+ top-5 at [MASK]:
74
  # 1 16.17 ' Paris'
75
  # 2 13.41 ' Strasbourg'
76
  # 3 13.35 'Paris'
 
78
  # 5 11.87 ' Versailles'
79
  ```
80
 
81
+ The same server also serves per-token embeddings directly:
82
 
83
  ```bash
 
84
  curl -s http://localhost:8080/embedding -d '{"content": "hello world"}'
85
  ```
86
 
fill-mask.py ADDED
@@ -0,0 +1,52 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # /// script
2
+ # requires-python = ">=3.10"
3
+ # dependencies = ["numpy", "requests", "gguf"]
4
+ # ///
5
+
6
+ # fill-mask.py — masked-token prediction against a stock llama-server.
7
+ #
8
+ # The encoder's MLM head is tied to the token embeddings, so the logits at the
9
+ # mask position are just `hidden @ token_embd^T`: fetch the per-token hidden
10
+ # states from `llama-server --embeddings --pooling none`, read the embedding
11
+ # matrix straight out of the GGUF, and take the top-K at the mask position.
12
+ #
13
+ # llama-server -m LFM2.5-Encoder-230M-F16.gguf --embeddings --pooling none
14
+ # uv run fill-mask.py LFM2.5-Encoder-230M-F16.gguf "The capital of France is [MASK]."
15
+ import sys
16
+
17
+ import numpy as np
18
+ import requests
19
+ from gguf import GGUFReader
20
+
21
+ gguf_path, prompt = sys.argv[1], sys.argv[2]
22
+ topk = int(sys.argv[3]) if len(sys.argv) > 3 else 5
23
+
24
+ # token_embd from the GGUF (memory-mapped; fp16/fp32 tensors read directly)
25
+ reader = GGUFReader(gguf_path)
26
+ embd = next(t for t in reader.tensors if t.name == "token_embd.weight")
27
+ W = np.array(embd.data).astype(np.float32) # [n_vocab, n_embd]
28
+
29
+ # tokenize server-side, replacing [MASK] with the model's mask token id
30
+ def tokenize(text: str, special: bool) -> list[int]:
31
+ r = requests.post("http://localhost:8080/tokenize",
32
+ json={"content": text, "add_special": special, "parse_special": True})
33
+ return r.json()["tokens"]
34
+
35
+ meta = {f.name: f for f in reader.fields.values()}
36
+ mask_id = int(meta["tokenizer.ggml.mask_token_id"].parts[-1][0])
37
+
38
+ pre, _, post = prompt.partition("[MASK]")
39
+ toks = tokenize(pre, True) + [mask_id] + tokenize(post, False)
40
+ pos = toks.index(mask_id)
41
+
42
+ # one non-causal forward; per-token hidden states
43
+ r = requests.post("http://localhost:8080/embedding",
44
+ json={"content": toks})
45
+ hidden = np.array(r.json()[0]["embedding"], dtype=np.float32) # [n_tok, n_embd]
46
+
47
+ logits = hidden[pos] @ W.T
48
+ top = np.argsort(logits)[::-1][:topk]
49
+ detok = lambda t: requests.post("http://localhost:8080/detokenize", json={"tokens": [int(t)]}).json()["content"]
50
+ print(f"top-{topk} at [MASK]:")
51
+ for i, t in enumerate(top, 1):
52
+ print(f" {i:>2} {logits[t]:9.4f} '{detok(t)}'")