anurag051194 commited on
Commit
675ceab
·
verified ·
1 Parent(s): a1e3de8

Rename Meridian-smaller to TwIL-LM3-Pro (model card and GGUF filenames)

Browse files
.gitattributes CHANGED
@@ -39,3 +39,8 @@ Meridian-smaller-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
39
  Meridian-smaller-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
40
  Meridian-smaller-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
41
  benchmarks.png filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
39
  Meridian-smaller-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
40
  Meridian-smaller-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
41
  benchmarks.png filter=lfs diff=lfs merge=lfs -text
42
+ TwIL-LM3-Pro-F16.gguf filter=lfs diff=lfs merge=lfs -text
43
+ TwIL-LM3-Pro-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
44
+ TwIL-LM3-Pro-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
45
+ TwIL-LM3-Pro-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
46
+ TwIL-LM3-Pro-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -6,7 +6,7 @@ pipeline_tag: text-generation
6
  base_model: ibm-granite/granite-4.2-3b
7
  license: other
8
  license_name: webai-non-commercial-license-ver.-1.0
9
- license_link: https://huggingface.co/webAI-Official/Meridian-smaller/blob/main/LICENSE.md
10
  tags:
11
  - granite
12
  - granite-4.2
@@ -21,7 +21,7 @@ tags:
21
  - meridian
22
  ---
23
 
24
- # Meridian-smaller
25
 
26
  A 3.66B reasoning model for **formal logic** tasks, built from
27
  [`ibm-granite/granite-4.2-3b`](https://huggingface.co/ibm-granite/granite-4.2-3b) through LoRA
@@ -36,7 +36,7 @@ gate, strict-7 and six-lane average of any arm in the tables below for which eac
36
  computed, including Qwen3-8B, Qwen3.5-4B, VibeThinker-3B and gpt-oss-120b (the 120B has no gate or strict-7
37
  value).
38
 
39
- ![Meridian-smaller formal and general reasoning benchmarks against VibeThinker-3B, Qwen3.5-4B, Qwen3-8B, gpt-oss-120b, LFM2.5-8B-A1B and TwIL-LM3](benchmarks.png)
40
 
41
  ## Highlights
42
 
@@ -58,11 +58,11 @@ value).
58
  Qwen3-8B (0.8493 / 0.7591) and gpt-oss-120b (0.8689 / 0.8086). It gains on BBH-logic
59
  (0.9013 → 0.9540), MATH-500 (0.6567 → 0.7467) and MuSR (0.5922 → 0.6409), and gives back
60
  GSM-Symbolic (0.8900 → 0.8267) and ARC (0.8933 → 0.8600).
61
- * **Against VibeThinker-3B, a reasoning model of similar size.** Meridian-smaller leads it on all
62
  six Track A lanes — macro gate 0.5539 against 0.4118, strict-7 0.2879 against 0.2021 — but not
63
  on Track B, where VibeThinker-3B is ahead on the 10-dataset macro (0.8097 against 0.7901) and
64
- Meridian-smaller is ahead on the 14-dataset macro (0.7425 against 0.7262). That 14-dataset lead
65
- comes entirely from BBH-logic (0.9540 against 0.6107); without that row Meridian-smaller trails.
66
  * **Structured formal output.** Tuned for the objects rather than the prose: FOL translation,
67
  entailment labels, semantic parses, Lean statements and Lean proof critique.
68
  * **Runs anywhere.** 3.66B parameters in bf16 (6.82 GiB), with a Q4\_K\_M GGUF at 2.09 GiB for
@@ -79,7 +79,7 @@ assistant — there is no safety or preference tuning here beyond what Granite 4
79
 
80
  | Property | Value |
81
  | ------------------------- | ---------------------------------------------------------------------------------------------- |
82
- | Model ID | `webAI-Official/Meridian-smaller` |
83
  | Base model | [`ibm-granite/granite-4.2-3b`](https://huggingface.co/ibm-granite/granite-4.2-3b) |
84
  | Total parameters | 3.66B (3,659,737,600) |
85
  | Architecture | Granite decoder-only dense transformer (`GraniteForCausalLM`); 40 layers, hidden size 2560, 40 attention heads / 8 KV heads |
@@ -102,7 +102,7 @@ here.
102
 
103
  ### Track A — in-domain formal logic
104
 
105
- Meridian-smaller, its base and VibeThinker-3B were run through the same harness, prompts, decoding
106
  settings and sampled rows described under [Evaluation protocol](#evaluation-protocol), and the
107
  other peer columns are the values already published for TwIL-LM3 on its card, produced by that
108
  same harness (see [Comparability](#limitations-and-caveats)). Throughput rows come from a
@@ -110,7 +110,7 @@ dedicated decode-throughput protocol: `ans/s` is defined throughout as
110
  `tok/s ÷ mean generation length`, so it measures completed answers rather than raw decode rate.
111
  Cells marked † need the engine note below.
112
 
113
- | lane / metric | Meridian-smaller | Granite-4.2-3B base | VibeThinker-3B | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
114
  |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
115
  | lean_formalize token_f1 | 0.5092 | 0.2943 | 0.2087 | 0.4996 | 0.5869 | 0.3690 | 0.1321 | 0.4655 | 0.4022 | **0.6306** |
116
  | rule_induction derivation | 0.4195 | 0.2267 | 0.2038 | 0.5078 | 0.3192 | 0.0825 | 0.0615 | 0.1936 | 0.3680 | **0.6518** |
@@ -137,7 +137,7 @@ average cannot be computed for it; that is what the — cells mean, not a zero.
137
  format and tokenizer make the corpus lanes score a different quantity. The number is reported
138
  for completeness but is not a comparable measurement, and is excluded from the bolding.
139
 
140
- † **Throughput for Meridian-smaller, its base and VibeThinker-3B** was measured with the same
141
  dedicated protocol and prompt file as the peer columns (128 prompts × 512 generated tokens, EOS
142
  ignored, greedy, `gpu_memory_utilization` 0.45, idle GPU), run on a single H200 with vLLM
143
  0.19.1 in a later session, whereas the other columns are figures recorded earlier on vLLM 0.11.2. Engine effects
@@ -147,7 +147,7 @@ against its published 15,880 (−4%), while VibeThinker-3B measured
147
  cells are therefore comparable with each other, and only approximately with the older columns.
148
  The Qwen3.5-4B figure comes from the earlier throughput sweep, which ran the Qwen3.5 models on the
149
  newer engine (vLLM 0.19.1) according to its driver script, so it belongs with the † cells.
150
- Meridian-smaller's two runs gave 20,454 and 21,884 tok/s (the first paid for cold
151
  kernel caches) and the table shows their mean. `mean gen length` is measured directly under the
152
  shared evaluation protocol and is comparable across every column.
153
 
@@ -165,7 +165,7 @@ equal-weight mean of five objectives: the four bounded classification lanes (`en
165
  derivation score. In the gate, `mcq_answer` and `procedural` are credited as
166
  `max(exact_match, loose_match)`: for free-text answer lanes, a response that is correct but
167
  differently formatted is a formatting artefact rather than a reasoning failure. This affects the
168
- aggregate only — the per-lane rows above stay strict. (Meridian-smaller's loose-match MCQ is
169
  0.6500 against the 0.4100 strict figure shown in the lane row.)
170
 
171
  **`macro_primary`** is the same mean over the four classification lanes alone, without
@@ -177,23 +177,23 @@ aggregate only — the per-lane rows above stay strict. (Meridian-smaller's loos
177
  harsh — exact match on generative lanes is near zero for every arm — so it is useful for ranking
178
  models against each other but not as an absolute capability measure.
179
 
180
- Meridian-smaller leads all four summary rows that every arm with a computable value can be
181
  compared on. The clearest margins are over the arms at its own scale and above: 0.5539 against
182
  0.3757 on the gate for LFM2.5-8B-A1B, and 0.2879 against 0.1971 on strict-7 for TwIL-LM3. Against
183
  Qwen3-8B the gate gap is only 0.020, but strict-7 is 0.2879 against 0.2093 — a difference that
184
- does not depend on loose-match credit — and Meridian-smaller leads on MCQ (strict 0.4100 against
185
  0.0000), rule induction (0.4195 against 0.3680) and both perplexity lanes.
186
 
187
  VibeThinker-3B — WeiboAI's 3B reasoning model, with 37.1% of its Track A generations truncated
188
- — trails Meridian-smaller on every objective lane and every summary row: macro gate 0.4118 against
189
  0.5539, strict-7 0.2021 against 0.2879, six-lane average 0.3291 against 0.5389. The gap is widest
190
  on Lean formalisation (0.2087 against 0.5092) and narrowest on entailment (0.5500 against 0.6700).
191
  Its corpus perplexities (16.93 and 27.03) are not comparable with the Granite-tokenizer models',
192
  because perplexity is per token and its vocabulary is 151,936 against 100,352.
193
 
194
  Qwen3.5-4B, a 4B reasoning model, sits between
195
- the small models and Meridian-smaller: macro gate 0.4466 against 0.5539, strict-7 0.1121 against
196
- 0.2879. It is ahead of Meridian-smaller on two rows only, `rule_induction` (0.5078 against 0.4195,
197
  second in the table after gpt-oss-120b) and `math_corpus` perplexity (3.5926 against 3.6983), and
198
  far behind on entailment (0.2400 against 0.6700) and strict MCQ (0.0000 against 0.4100).
199
 
@@ -205,7 +205,7 @@ spots in absolute terms are `procedural` (strict 0.1200, loose 0.2350) and FOL t
205
 
206
  ### Track B — held-out benchmarks
207
 
208
- | dataset | Meridian-smaller | Granite-4.2-3B base | VibeThinker-3B | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
209
  |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
210
  | gsm8k | 0.9433 | 0.9533 | 0.9600 | 0.8633 | 0.8733 | 0.8300 | 0.8767 | 0.9133 | 0.9567 | **0.9767** |
211
  | svamp | **0.9500** | 0.9200 | 0.9367 | 0.8867 | 0.8500 | 0.8200 | 0.9000 | 0.9133 | 0.9400 | 0.9400 |
@@ -230,13 +230,13 @@ spots in absolute terms are `procedural` (strict 0.1200, loose 0.2350) and FOL t
230
  ‡ MXFP4 weights, tensor-parallel 2 — quantized and multi-GPU, so not directly comparable to the
231
  single-GPU BF16 rows. ¶ 74% of its `rudas_ood` generations hit the length cap, so that cell is a
232
  truncation artefact rather than a measured score; excluding the row, its 13-dataset macro is
233
- 0.8708. ¶¶ 93.7% of Meridian-smaller's `rudas_ood` generations also hit the length cap, so its
234
  cell is likewise a truncation artefact rather than a measurement of the model's ability.
235
  ¶¶¶ Qwen3.5-4B hits the length cap on 67.3% of IFEval, 59.0% of MATH-500 and 99.3% of `rudas_ood`
236
  generations, so those three cells are truncation artefacts too.
237
 
238
  Lengths marked ≈ are derived from stored generations rather than read from the run. For
239
- Meridian-smaller, its base and VibeThinker-3B they were re-tokenized directly with each model's
240
  own tokenizer and
241
  averaged over the 14 datasets (MuSR counted once); the same method reproduces TwIL-LM3's measured
242
  482 to within 1%. The peer lengths marked ≈ use each model's characters-per-token ratio and are
@@ -245,7 +245,7 @@ the engine note under that table) and `ans/s` divides it by the mean generation
245
  VibeThinker-3B Track B rows come from the same external-baseline run and vLLM 0.11.2 engine as
246
  the other external peers.
247
 
248
- The honest summary of this table is that Meridian-smaller does not lead it. Larger models score
249
  higher, and gpt-oss-120b leads seven of the fourteen dataset rows. Three things are worth
250
  extracting anyway. First, it holds its own base on the held-out suite (10-dataset macro 0.7901
251
  against 0.7942, a difference well inside the sampling noise at n = 300 per dataset) while gaining
@@ -254,14 +254,14 @@ over the base, driven by BBH-logic, MATH-500 and MuSR. Third, it edges LFM2.5-8B
254
  macros at less than half the parameters and leads the table outright on SVAMP (0.9500).
255
 
256
  VibeThinker-3B is the stronger held-out model on the 10-dataset macro (0.8097 against 0.7901). It
257
- is ahead of Meridian-smaller on seven datasets — GSM8K, ARC, LogicBench, StrategyQA, DROP,
258
  MMLU-Redux and MATH-500 — and behind on the other seven: SVAMP, GSM-Symbolic, CSQA, MuSR, IFEval,
259
- BBH-logic and `rudas_ood`. Meridian-smaller's 0.0163 lead on the 14-dataset macro comes entirely
260
  from BBH-logic (0.9540 against 0.6107): on the other thirteen datasets it averages 0.7263 against
261
  VibeThinker-3B's 0.7350. VibeThinker-3B also writes much longer answers on Track B (about 1,789
262
  tokens against 792).
263
 
264
- Track B here was run with the chat template's thinking mode **disabled** for Meridian-smaller and
265
  its base (the prompt ends in an empty `<think></think>`), as it was for TwIL-LM3, whereas Track A
266
  uses the default thinking mode. The Track B numbers therefore describe non-reasoning behaviour;
267
  they are not a measure of what a thinking-mode generation would score. Some Track B cells are also
@@ -272,12 +272,12 @@ MuSR-team sit slightly above the 2% cap-hit threshold (3.0%, 3.0% and 2.8%).
272
 
273
  The tables above use the public VibeThinker-3B checkpoint. The same post-training pipeline was also
274
  applied to it, and the best tuned configuration (WiSE-FT, λ = 0.50) is the closest same-scale
275
- comparison to Meridian-smaller. These values come from the family comparison tables rather than
276
  from a per-lane raw report, so they are shown as a summary only:
277
 
278
  | model | macro gate | macro_primary | B10 | B14 | Track A truncation |
279
  |---|---:|---:|---:|---:|---:|
280
- | Meridian-smaller | **0.554** | **0.588** | 0.790 | **0.743** | 24.2% |
281
  | VibeThinker-3B, WiSE-FT λ = 0.50 | 0.541 | **0.588** | **0.802** | 0.728 ◊ | 14.3% |
282
 
283
  ◊ There is no B14 row for the λ = 0.50 configuration; the figure is the best tuned VibeThinker-3B
@@ -285,7 +285,7 @@ B14 in the family table (SLERP, d = 0.5, t = 0.5).
285
 
286
  The two are effectively tied on Track A (macro gate 0.554 against 0.541, `macro_primary` equal at
287
  0.588, both within sampling noise at n = 200), and the tuned VibeThinker-3B is ahead on the
288
- 10-dataset macro (0.802 against 0.790) with a lower truncation rate. Meridian-smaller's edge is the
289
  14-dataset macro (0.743 against 0.728), a gap that cannot be broken down per dataset from the
290
  summary values. Read gaps to the
291
  *untuned* VibeThinker-3B within one source: the family table records that checkpoint at a gate of
@@ -319,7 +319,7 @@ is a selection tool, not a benchmark claim.
319
  import torch
320
  from transformers import AutoModelForCausalLM, AutoTokenizer
321
 
322
- model_id = "webAI-Official/Meridian-smaller"
323
  tok = AutoTokenizer.from_pretrained(model_id)
324
  model = AutoModelForCausalLM.from_pretrained(
325
  model_id, torch_dtype=torch.bfloat16, device_map="auto"
@@ -357,14 +357,14 @@ as EOS, so run it in conversation mode (`-cnv`).
357
 
358
  | file | quant | size | bits/weight | notes |
359
  |---|---|---:|---:|---|
360
- | `Meridian-smaller-Q4_K_M.gguf` | Q4_K_M | 2.09 GiB | 4.91 | recommended default; runs on CPU or 4 GB of VRAM |
361
- | `Meridian-smaller-Q5_K_M.gguf` | Q5_K_M | 2.43 GiB | 5.71 | a little more headroom than Q4_K_M |
362
- | `Meridian-smaller-Q6_K.gguf` | Q6_K | 2.80 GiB | 6.57 | close to Q8_0 quality at about three-quarters the size |
363
- | `Meridian-smaller-Q8_0.gguf` | Q8_0 | 3.63 GiB | 8.51 | near-lossless, for quality-sensitive use |
364
- | `Meridian-smaller-F16.gguf` | F16 | 6.82 GiB | 16.01 | unquantized, for requantization or reference runs |
365
 
366
  ```bash
367
- llama-cli -hf webAI-Official/Meridian-smaller:Q4_K_M -cnv --temp 0 -n 2048
368
  ```
369
 
370
  Two things matter for reproducing the scores above under llama.cpp. Pass `--temp 0`, because the
@@ -413,12 +413,12 @@ makes no claim about those. The weak absolute areas inside the specialisation ar
413
  (exact match 0.0100), `procedural` (strict 0.1200) and semantic parsing exact match (0.0000);
414
  `rule_induction` parses only 56.5% of outputs.
415
 
416
- **Comparability.** For Track A, Meridian-smaller, its base, VibeThinker-3B, Qwen3.5-4B, Qwen3-8B, LFM2-2.6B,
417
  LFM2.5-8B-A1B and Llama-3.2-3B were checked to share the same sampled-row manifest and dataset
418
  hash, seed and decoding; the TwIL-LM3 and gpt-oss-120b values are carried over from the TwIL-LM3
419
  card, which describes the same harness. For Track B, the arms checked (including VibeThinker-3B
420
  and Qwen3.5-4B, on all 18 tasks) share the same sampled rows and decoding, but the serving engine differs between
421
- arms (vLLM 0.19.1 for Meridian-smaller, its base, Qwen3.5-4B, Qwen3-8B and LFM2.5-8B-A1B; vLLM 0.11.2 for
422
  TwIL-LM3, Llama-3.2-3B and VibeThinker-3B), and the engine version is part of the protocol hash.
423
  Throughput has its own, separate engine caveat (see the † note under the Track A table). With
424
  n = 200 per lane on Track A and n = 300 per dataset on Track B, differences of two to three points
@@ -435,7 +435,7 @@ are within sampling noise.
435
  sampled rows.
436
  - Throughput: 128 prompts drawn from a fixed Track A prompt file, 512 generated tokens each with
437
  EOS ignored, greedy, vLLM `gpu_memory_utilization` 0.45, `max_model_len` 4096, on an otherwise
438
- idle GPU (a single H200 for the Meridian-smaller, base and VibeThinker-3B runs); the reported rate is generated tokens over decode time, excluding engine start-up and
439
  compilation.
440
 
441
  `repetition_penalty = 1.0` is load-bearing. A 1.1 penalty produced apparent 20-point swings on
@@ -444,7 +444,7 @@ identity so a mismatched runner fails loudly instead of quietly producing a diff
444
 
445
  ## Relationship to TwIL-LM
446
 
447
- Meridian-smaller applies the same post-training pipeline as the
448
  [TwIL-LM3](https://huggingface.co/webAI-Official/TwIL-LM3) and TwIL-LM family — LoRA SFT,
449
  checkpoint fusion, WiSE-FT and MGPO — to a different base, IBM's Granite 4.2 3B, instead of
450
  SmolLM3 or SmolLM2 with some additional mechanisms. Compared with TwIL-LM3 it is a stronger in-domain model (macro gate 0.5539
 
6
  base_model: ibm-granite/granite-4.2-3b
7
  license: other
8
  license_name: webai-non-commercial-license-ver.-1.0
9
+ license_link: https://huggingface.co/webAI-Official/TwIL-LM3-Pro/blob/main/LICENSE.md
10
  tags:
11
  - granite
12
  - granite-4.2
 
21
  - meridian
22
  ---
23
 
24
+ # TwIL-LM3-Pro
25
 
26
  A 3.66B reasoning model for **formal logic** tasks, built from
27
  [`ibm-granite/granite-4.2-3b`](https://huggingface.co/ibm-granite/granite-4.2-3b) through LoRA
 
36
  computed, including Qwen3-8B, Qwen3.5-4B, VibeThinker-3B and gpt-oss-120b (the 120B has no gate or strict-7
37
  value).
38
 
39
+ ![TwIL-LM3-Pro formal and general reasoning benchmarks against VibeThinker-3B, Qwen3.5-4B, Qwen3-8B, gpt-oss-120b, LFM2.5-8B-A1B and TwIL-LM3](benchmarks.png)
40
 
41
  ## Highlights
42
 
 
58
  Qwen3-8B (0.8493 / 0.7591) and gpt-oss-120b (0.8689 / 0.8086). It gains on BBH-logic
59
  (0.9013 → 0.9540), MATH-500 (0.6567 → 0.7467) and MuSR (0.5922 → 0.6409), and gives back
60
  GSM-Symbolic (0.8900 → 0.8267) and ARC (0.8933 → 0.8600).
61
+ * **Against VibeThinker-3B, a reasoning model of similar size.** TwIL-LM3-Pro leads it on all
62
  six Track A lanes — macro gate 0.5539 against 0.4118, strict-7 0.2879 against 0.2021 — but not
63
  on Track B, where VibeThinker-3B is ahead on the 10-dataset macro (0.8097 against 0.7901) and
64
+ TwIL-LM3-Pro is ahead on the 14-dataset macro (0.7425 against 0.7262). That 14-dataset lead
65
+ comes entirely from BBH-logic (0.9540 against 0.6107); without that row TwIL-LM3-Pro trails.
66
  * **Structured formal output.** Tuned for the objects rather than the prose: FOL translation,
67
  entailment labels, semantic parses, Lean statements and Lean proof critique.
68
  * **Runs anywhere.** 3.66B parameters in bf16 (6.82 GiB), with a Q4\_K\_M GGUF at 2.09 GiB for
 
79
 
80
  | Property | Value |
81
  | ------------------------- | ---------------------------------------------------------------------------------------------- |
82
+ | Model ID | `webAI-Official/TwIL-LM3-Pro` |
83
  | Base model | [`ibm-granite/granite-4.2-3b`](https://huggingface.co/ibm-granite/granite-4.2-3b) |
84
  | Total parameters | 3.66B (3,659,737,600) |
85
  | Architecture | Granite decoder-only dense transformer (`GraniteForCausalLM`); 40 layers, hidden size 2560, 40 attention heads / 8 KV heads |
 
102
 
103
  ### Track A — in-domain formal logic
104
 
105
+ TwIL-LM3-Pro, its base and VibeThinker-3B were run through the same harness, prompts, decoding
106
  settings and sampled rows described under [Evaluation protocol](#evaluation-protocol), and the
107
  other peer columns are the values already published for TwIL-LM3 on its card, produced by that
108
  same harness (see [Comparability](#limitations-and-caveats)). Throughput rows come from a
 
110
  `tok/s ÷ mean generation length`, so it measures completed answers rather than raw decode rate.
111
  Cells marked † need the engine note below.
112
 
113
+ | lane / metric | TwIL-LM3-Pro | Granite-4.2-3B base | VibeThinker-3B | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
114
  |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
115
  | lean_formalize token_f1 | 0.5092 | 0.2943 | 0.2087 | 0.4996 | 0.5869 | 0.3690 | 0.1321 | 0.4655 | 0.4022 | **0.6306** |
116
  | rule_induction derivation | 0.4195 | 0.2267 | 0.2038 | 0.5078 | 0.3192 | 0.0825 | 0.0615 | 0.1936 | 0.3680 | **0.6518** |
 
137
  format and tokenizer make the corpus lanes score a different quantity. The number is reported
138
  for completeness but is not a comparable measurement, and is excluded from the bolding.
139
 
140
+ † **Throughput for TwIL-LM3-Pro, its base and VibeThinker-3B** was measured with the same
141
  dedicated protocol and prompt file as the peer columns (128 prompts × 512 generated tokens, EOS
142
  ignored, greedy, `gpu_memory_utilization` 0.45, idle GPU), run on a single H200 with vLLM
143
  0.19.1 in a later session, whereas the other columns are figures recorded earlier on vLLM 0.11.2. Engine effects
 
147
  cells are therefore comparable with each other, and only approximately with the older columns.
148
  The Qwen3.5-4B figure comes from the earlier throughput sweep, which ran the Qwen3.5 models on the
149
  newer engine (vLLM 0.19.1) according to its driver script, so it belongs with the † cells.
150
+ TwIL-LM3-Pro's two runs gave 20,454 and 21,884 tok/s (the first paid for cold
151
  kernel caches) and the table shows their mean. `mean gen length` is measured directly under the
152
  shared evaluation protocol and is comparable across every column.
153
 
 
165
  derivation score. In the gate, `mcq_answer` and `procedural` are credited as
166
  `max(exact_match, loose_match)`: for free-text answer lanes, a response that is correct but
167
  differently formatted is a formatting artefact rather than a reasoning failure. This affects the
168
+ aggregate only — the per-lane rows above stay strict. (TwIL-LM3-Pro's loose-match MCQ is
169
  0.6500 against the 0.4100 strict figure shown in the lane row.)
170
 
171
  **`macro_primary`** is the same mean over the four classification lanes alone, without
 
177
  harsh — exact match on generative lanes is near zero for every arm — so it is useful for ranking
178
  models against each other but not as an absolute capability measure.
179
 
180
+ TwIL-LM3-Pro leads all four summary rows that every arm with a computable value can be
181
  compared on. The clearest margins are over the arms at its own scale and above: 0.5539 against
182
  0.3757 on the gate for LFM2.5-8B-A1B, and 0.2879 against 0.1971 on strict-7 for TwIL-LM3. Against
183
  Qwen3-8B the gate gap is only 0.020, but strict-7 is 0.2879 against 0.2093 — a difference that
184
+ does not depend on loose-match credit — and TwIL-LM3-Pro leads on MCQ (strict 0.4100 against
185
  0.0000), rule induction (0.4195 against 0.3680) and both perplexity lanes.
186
 
187
  VibeThinker-3B — WeiboAI's 3B reasoning model, with 37.1% of its Track A generations truncated
188
+ — trails TwIL-LM3-Pro on every objective lane and every summary row: macro gate 0.4118 against
189
  0.5539, strict-7 0.2021 against 0.2879, six-lane average 0.3291 against 0.5389. The gap is widest
190
  on Lean formalisation (0.2087 against 0.5092) and narrowest on entailment (0.5500 against 0.6700).
191
  Its corpus perplexities (16.93 and 27.03) are not comparable with the Granite-tokenizer models',
192
  because perplexity is per token and its vocabulary is 151,936 against 100,352.
193
 
194
  Qwen3.5-4B, a 4B reasoning model, sits between
195
+ the small models and TwIL-LM3-Pro: macro gate 0.4466 against 0.5539, strict-7 0.1121 against
196
+ 0.2879. It is ahead of TwIL-LM3-Pro on two rows only, `rule_induction` (0.5078 against 0.4195,
197
  second in the table after gpt-oss-120b) and `math_corpus` perplexity (3.5926 against 3.6983), and
198
  far behind on entailment (0.2400 against 0.6700) and strict MCQ (0.0000 against 0.4100).
199
 
 
205
 
206
  ### Track B — held-out benchmarks
207
 
208
+ | dataset | TwIL-LM3-Pro | Granite-4.2-3B base | VibeThinker-3B | Qwen3.5-4B | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
209
  |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
210
  | gsm8k | 0.9433 | 0.9533 | 0.9600 | 0.8633 | 0.8733 | 0.8300 | 0.8767 | 0.9133 | 0.9567 | **0.9767** |
211
  | svamp | **0.9500** | 0.9200 | 0.9367 | 0.8867 | 0.8500 | 0.8200 | 0.9000 | 0.9133 | 0.9400 | 0.9400 |
 
230
  ‡ MXFP4 weights, tensor-parallel 2 — quantized and multi-GPU, so not directly comparable to the
231
  single-GPU BF16 rows. ¶ 74% of its `rudas_ood` generations hit the length cap, so that cell is a
232
  truncation artefact rather than a measured score; excluding the row, its 13-dataset macro is
233
+ 0.8708. ¶¶ 93.7% of TwIL-LM3-Pro's `rudas_ood` generations also hit the length cap, so its
234
  cell is likewise a truncation artefact rather than a measurement of the model's ability.
235
  ¶¶¶ Qwen3.5-4B hits the length cap on 67.3% of IFEval, 59.0% of MATH-500 and 99.3% of `rudas_ood`
236
  generations, so those three cells are truncation artefacts too.
237
 
238
  Lengths marked ≈ are derived from stored generations rather than read from the run. For
239
+ TwIL-LM3-Pro, its base and VibeThinker-3B they were re-tokenized directly with each model's
240
  own tokenizer and
241
  averaged over the 14 datasets (MuSR counted once); the same method reproduces TwIL-LM3's measured
242
  482 to within 1%. The peer lengths marked ≈ use each model's characters-per-token ratio and are
 
245
  VibeThinker-3B Track B rows come from the same external-baseline run and vLLM 0.11.2 engine as
246
  the other external peers.
247
 
248
+ The honest summary of this table is that TwIL-LM3-Pro does not lead it. Larger models score
249
  higher, and gpt-oss-120b leads seven of the fourteen dataset rows. Three things are worth
250
  extracting anyway. First, it holds its own base on the held-out suite (10-dataset macro 0.7901
251
  against 0.7942, a difference well inside the sampling noise at n = 300 per dataset) while gaining
 
254
  macros at less than half the parameters and leads the table outright on SVAMP (0.9500).
255
 
256
  VibeThinker-3B is the stronger held-out model on the 10-dataset macro (0.8097 against 0.7901). It
257
+ is ahead of TwIL-LM3-Pro on seven datasets — GSM8K, ARC, LogicBench, StrategyQA, DROP,
258
  MMLU-Redux and MATH-500 — and behind on the other seven: SVAMP, GSM-Symbolic, CSQA, MuSR, IFEval,
259
+ BBH-logic and `rudas_ood`. TwIL-LM3-Pro's 0.0163 lead on the 14-dataset macro comes entirely
260
  from BBH-logic (0.9540 against 0.6107): on the other thirteen datasets it averages 0.7263 against
261
  VibeThinker-3B's 0.7350. VibeThinker-3B also writes much longer answers on Track B (about 1,789
262
  tokens against 792).
263
 
264
+ Track B here was run with the chat template's thinking mode **disabled** for TwIL-LM3-Pro and
265
  its base (the prompt ends in an empty `<think></think>`), as it was for TwIL-LM3, whereas Track A
266
  uses the default thinking mode. The Track B numbers therefore describe non-reasoning behaviour;
267
  they are not a measure of what a thinking-mode generation would score. Some Track B cells are also
 
272
 
273
  The tables above use the public VibeThinker-3B checkpoint. The same post-training pipeline was also
274
  applied to it, and the best tuned configuration (WiSE-FT, λ = 0.50) is the closest same-scale
275
+ comparison to TwIL-LM3-Pro. These values come from the family comparison tables rather than
276
  from a per-lane raw report, so they are shown as a summary only:
277
 
278
  | model | macro gate | macro_primary | B10 | B14 | Track A truncation |
279
  |---|---:|---:|---:|---:|---:|
280
+ | TwIL-LM3-Pro | **0.554** | **0.588** | 0.790 | **0.743** | 24.2% |
281
  | VibeThinker-3B, WiSE-FT λ = 0.50 | 0.541 | **0.588** | **0.802** | 0.728 ◊ | 14.3% |
282
 
283
  ◊ There is no B14 row for the λ = 0.50 configuration; the figure is the best tuned VibeThinker-3B
 
285
 
286
  The two are effectively tied on Track A (macro gate 0.554 against 0.541, `macro_primary` equal at
287
  0.588, both within sampling noise at n = 200), and the tuned VibeThinker-3B is ahead on the
288
+ 10-dataset macro (0.802 against 0.790) with a lower truncation rate. TwIL-LM3-Pro's edge is the
289
  14-dataset macro (0.743 against 0.728), a gap that cannot be broken down per dataset from the
290
  summary values. Read gaps to the
291
  *untuned* VibeThinker-3B within one source: the family table records that checkpoint at a gate of
 
319
  import torch
320
  from transformers import AutoModelForCausalLM, AutoTokenizer
321
 
322
+ model_id = "webAI-Official/TwIL-LM3-Pro"
323
  tok = AutoTokenizer.from_pretrained(model_id)
324
  model = AutoModelForCausalLM.from_pretrained(
325
  model_id, torch_dtype=torch.bfloat16, device_map="auto"
 
357
 
358
  | file | quant | size | bits/weight | notes |
359
  |---|---|---:|---:|---|
360
+ | `TwIL-LM3-Pro-Q4_K_M.gguf` | Q4_K_M | 2.09 GiB | 4.91 | recommended default; runs on CPU or 4 GB of VRAM |
361
+ | `TwIL-LM3-Pro-Q5_K_M.gguf` | Q5_K_M | 2.43 GiB | 5.71 | a little more headroom than Q4_K_M |
362
+ | `TwIL-LM3-Pro-Q6_K.gguf` | Q6_K | 2.80 GiB | 6.57 | close to Q8_0 quality at about three-quarters the size |
363
+ | `TwIL-LM3-Pro-Q8_0.gguf` | Q8_0 | 3.63 GiB | 8.51 | near-lossless, for quality-sensitive use |
364
+ | `TwIL-LM3-Pro-F16.gguf` | F16 | 6.82 GiB | 16.01 | unquantized, for requantization or reference runs |
365
 
366
  ```bash
367
+ llama-cli -hf webAI-Official/TwIL-LM3-Pro:Q4_K_M -cnv --temp 0 -n 2048
368
  ```
369
 
370
  Two things matter for reproducing the scores above under llama.cpp. Pass `--temp 0`, because the
 
413
  (exact match 0.0100), `procedural` (strict 0.1200) and semantic parsing exact match (0.0000);
414
  `rule_induction` parses only 56.5% of outputs.
415
 
416
+ **Comparability.** For Track A, TwIL-LM3-Pro, its base, VibeThinker-3B, Qwen3.5-4B, Qwen3-8B, LFM2-2.6B,
417
  LFM2.5-8B-A1B and Llama-3.2-3B were checked to share the same sampled-row manifest and dataset
418
  hash, seed and decoding; the TwIL-LM3 and gpt-oss-120b values are carried over from the TwIL-LM3
419
  card, which describes the same harness. For Track B, the arms checked (including VibeThinker-3B
420
  and Qwen3.5-4B, on all 18 tasks) share the same sampled rows and decoding, but the serving engine differs between
421
+ arms (vLLM 0.19.1 for TwIL-LM3-Pro, its base, Qwen3.5-4B, Qwen3-8B and LFM2.5-8B-A1B; vLLM 0.11.2 for
422
  TwIL-LM3, Llama-3.2-3B and VibeThinker-3B), and the engine version is part of the protocol hash.
423
  Throughput has its own, separate engine caveat (see the † note under the Track A table). With
424
  n = 200 per lane on Track A and n = 300 per dataset on Track B, differences of two to three points
 
435
  sampled rows.
436
  - Throughput: 128 prompts drawn from a fixed Track A prompt file, 512 generated tokens each with
437
  EOS ignored, greedy, vLLM `gpu_memory_utilization` 0.45, `max_model_len` 4096, on an otherwise
438
+ idle GPU (a single H200 for the TwIL-LM3-Pro, base and VibeThinker-3B runs); the reported rate is generated tokens over decode time, excluding engine start-up and
439
  compilation.
440
 
441
  `repetition_penalty = 1.0` is load-bearing. A 1.1 penalty produced apparent 20-point swings on
 
444
 
445
  ## Relationship to TwIL-LM
446
 
447
+ TwIL-LM3-Pro applies the same post-training pipeline as the
448
  [TwIL-LM3](https://huggingface.co/webAI-Official/TwIL-LM3) and TwIL-LM family — LoRA SFT,
449
  checkpoint fusion, WiSE-FT and MGPO — to a different base, IBM's Granite 4.2 3B, instead of
450
  SmolLM3 or SmolLM2 with some additional mechanisms. Compared with TwIL-LM3 it is a stronger in-domain model (macro gate 0.5539
Meridian-smaller-F16.gguf → TwIL-LM3-Pro-F16.gguf RENAMED
File without changes
Meridian-smaller-Q4_K_M.gguf → TwIL-LM3-Pro-Q4_K_M.gguf RENAMED
File without changes
Meridian-smaller-Q5_K_M.gguf → TwIL-LM3-Pro-Q5_K_M.gguf RENAMED
File without changes
Meridian-smaller-Q6_K.gguf → TwIL-LM3-Pro-Q6_K.gguf RENAMED
File without changes
Meridian-smaller-Q8_0.gguf → TwIL-LM3-Pro-Q8_0.gguf RENAMED
File without changes