File size: 23,993 Bytes
0b79309
 
 
 
 
 
 
 
1ea6dac
0b79309
1ea6dac
 
 
 
 
3852c07
0b79309
 
 
 
 
 
 
 
 
 
 
 
 
 
3852c07
 
 
 
 
 
 
 
 
 
0b79309
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3852c07
 
 
 
 
 
 
 
 
 
0b79309
b1208f0
 
 
3852c07
 
11b6f0a
 
 
 
 
b1208f0
11b6f0a
 
b1208f0
 
 
 
 
 
11b6f0a
b1208f0
3852c07
76af275
 
81c08e1
 
 
 
 
 
 
fd5f34e
76af275
fd5f34e
76af275
 
 
fd5f34e
b1208f0
 
 
 
76af275
fd5f34e
76af275
 
b1208f0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fd5f34e
76af275
fd5f34e
b1208f0
 
 
 
fd5f34e
76af275
 
 
b1208f0
 
fd5f34e
76af275
b1208f0
 
fd5f34e
76af275
b1208f0
 
 
 
5b107df
76af275
5b107df
b1208f0
 
 
 
 
5b107df
76af275
5b107df
 
 
 
76af275
5b107df
 
 
 
76af275
5b107df
 
 
 
 
 
b1208f0
 
11b6f0a
b1208f0
 
 
 
11b6f0a
 
b1208f0
 
76af275
 
b1208f0
 
 
 
 
 
 
3852c07
76af275
 
b1208f0
 
 
 
 
d6a89c2
b1208f0
76af275
b1208f0
 
 
 
 
76af275
b1208f0
76af275
b1208f0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
76af275
 
b1208f0
 
76af275
b1208f0
 
 
 
81c08e1
b1208f0
 
 
 
 
 
76af275
81c08e1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b1208f0
 
 
 
81c08e1
 
b1208f0
d6a89c2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
81c08e1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b1208f0
 
 
81c08e1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b1208f0
 
 
 
 
81c08e1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3852c07
b1208f0
3852c07
b1208f0
3852c07
b1208f0
 
 
 
 
 
 
3852c07
b1208f0
 
5b107df
b1208f0
5b107df
 
b1208f0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5b107df
424cfe8
 
 
b1208f0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1787ffa
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
---
license: mit
language:
- en
pipeline_tag: text-generation
tags:
- function-calling
- tool-calling
- constrained-decoding
- grammar-constrained-decoding
- structured-generation
- json
- small-model
- slm
- on-device
- edge
library_name: pytorch
model-index:
- name: thimble-v6
  results:
  - task:
      type: text-generation
      name: Function calling (ordered strict exact match)
    dataset:
      name: Seal-Tools in-domain
      type: seal-tools
    metrics:
    - type: exact_match
      value: 33.1
      name: Seal-Tools in-domain
  - task:
      type: text-generation
      name: Function calling (ordered strict exact match)
    dataset:
      name: Seal-Tools out-of-domain
      type: seal-tools
    metrics:
    - type: exact_match
      value: 28.1
      name: Seal-Tools out-of-domain
  - task:
      type: text-generation
      name: Function calling (ordered strict exact match)
    dataset:
      name: Mobile Actions
      type: mobile-actions
    metrics:
    - type: exact_match
      value: 86.3
      name: Mobile Actions
  - task:
      type: text-generation
      name: Function calling (ordered strict exact match)
    dataset:
      name: DroidCall
      type: droidcall
    metrics:
    - type: exact_match
      value: 52.5
      name: DroidCall
  - task:
      type: text-generation
      name: Function calling (ordered strict exact match)
    dataset:
      name: BFCL v4 single-turn
      type: bfcl
    metrics:
    - type: exact_match
      value: 23.5
      name: BFCL v4 single-turn
---

<div align="center">

# 🧡 Thimble

### A tool-calling layer, not a language model.

**Your schemas in, validated calls out, at 48M parameters.**

[![License: MIT](https://img.shields.io/badge/license-MIT-1f6feb?style=flat-square)](https://github.com/nikshepsvn/thimble/blob/master/LICENSE)
[![Model on Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20model-thimble--v6-ffcc4d?style=flat-square)](https://huggingface.co/flashvenom/thimble)
[![Parameters](https://img.shields.io/badge/params-48.12M-c8324c?style=flat-square)](#numbers)
[![Well-formed JSON](https://img.shields.io/badge/well--formed%20JSON-100%25%20by%20construction-2ea043?style=flat-square)](#the-contract)
[![Build cost](https://img.shields.io/badge/total%20build%20cost-%24260-8957e5?style=flat-square)](https://github.com/nikshepsvn/thimble/blob/master/REPRODUCING.md)

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/nikshepsvn/thimble/master/assets/results-dark.png">
  <img alt="Accuracy by suite, split by whether the tool catalog appeared in training" src="https://raw.githubusercontent.com/nikshepsvn/thimble/master/assets/results-light.png" width="100%">
</picture>

</div>

It does not converse, reason, or write prose β€” it was never trained to. It reads
a catalog of typed functions and a request, and returns the calls to make or an
empty list when nothing fits.

That narrowness is the design, not a limitation of it. The tokenizer, the
training loss, and the decoder are all built around the same five decisions, so
the model is never asked to spend capacity on JSON it will never emit. The whole
job then fits in 48M parameters β€” small enough that specializing it to one API
surface is routine rather than a project.

## The contract

Three guarantees hold on **any** catalog, with no training and no configuration,
because they come from a grammar compiled out of your schemas rather than from
the weights:

- **Output is always well-formed JSON.** Not usually β€” always. Malformed output
  is unreachable, not unlikely.
- **Argument keys come from your schema.** Parameter-name hallucination is
  structurally impossible.
- **Calls to tools you did not declare cannot be emitted.**

The model is consulted at exactly five choice points: refuse-or-call, which tool,
include this optional, what value, stop or continue. Everything else β€” braces,
quotes, commas, every argument key β€” is determined before it runs.

Accuracy is a separate question, answered below with numbers. The contract is not
conditional on any of them.

## Try it in your browser

**[nikshepsvn.com/thimble](https://nikshepsvn.com/thimble/)** β€” the C engine
compiled to WebAssembly. The whole model runs in the tab (105KB engine + 48MB
int8 weights, no server); edit the tool catalog live and watch the grammar
adapt with zero retraining. ~250–650ms per call via SIMD128.

## Try it in 30 seconds

```
git clone https://github.com/nikshepsvn/thimble && cd thimble
uv venv && uv pip install -e ".[hub]"

# both files come from the HF repo; neither is in git
hf download flashvenom/thimble thimble-v6.pt --local-dir checkpoints/
hf download flashvenom/thimble tokenizer.json --local-dir data/
```

Then:

```
$ python demo.py "make a reservation at Nobu for 2 people at 7pm and text Sam saying dinner is on"
[
  {"name": "createReservation",
   "arguments": {"partySize": 2, "restaurant": "Nobu", "time": "7pm"}},
  {"name": "sendMessage",
   "arguments": {"body": "dinner is on", "contact": "Sam"}}
]

$ python demo.py "sing me a happy birthday song"
[]  (refused: no tool applies)
```

Real output, not a mock β€” the typed integer `partySize`, the two-call
composition, and the refusal. Point it at your own tools with:

```
python scripts/eval_catalog.py --ckpt thimble-v6 \
    --catalog my_tools.json --gold my_eval.jsonl
```

## Does this fit your problem?

**It works out of the box when** requests are command-shaped and state their
values: `annotate variant rs4988235 against build GRCh38`. Identifiers, codes,
dates, numbers, enum picks β€” copied, not inferred. Chains are fine: two-plus-call
rows score **73.5%** on a catalog it knows.

The gate is how *extractive* the request is, not what domain it belongs to. In a
pair of small probes, an unseen biomedical catalog in `dot.notation` scored 0.75
while a familiar-looking app catalog with conversational phrasing scored 0.57.
Unfamiliar vocabulary is survivable; phrasing that hides the values is not.
(Two hand-written probes, 15 rows β€” directional, not a measurement.)

**Adapt it when** you need conversational phrasing, disciplined handling of
optional arguments, or calibrated refusal. Those three are what specializing
buys, and they are the documented weak spots β€” see below.

**Use something else when** you have an open-world catalog, need Java or
JavaScript schema dialects, or need parallel instantiations of one schema. And
if you can afford 600M parameters, fine-tune Qwen instead β€” it will probably
score higher. This is for when you cannot: a memory ceiling, a latency floor, or
wanting a separate model per customer rather than one prompted model for all.

## Numbers

Ordered strict exact match β€” a row passes only if the function names, the call
order, and every argument value match. The right-hand column is a yardstick, not
a rival: Needle 2 (Cactus Compute, 45M params, 153B training tokens), their
published numbers on their metric, unmodified. It is there so the left column
has a scale.

**Catalog represented in training** (eval rows firewalled out):

| suite | Thimble v6 | Needle 2 (45M) |
|---|---:|---:|
| Mobile Actions (961) | **86.3** | 63.7 |
| Mobile Actions, two-plus-call rows | **73.5** | 48.4 |
| DroidCall (200) | **52.5** | 17.0 |
| Seal-Tools in-domain (700) | **33.1** | 32.6 |
| Well-formed JSON | **100.0** | 93.4 |

**Catalog never seen:**

| suite | Thimble v6 | Needle 2 (45M) |
|---|---:|---:|
| Seal-Tools out-of-domain (654) | 28.1 | 28.7 |
| BFCL v4 single-turn (3,641) | 23.5 | 42.6 |

The spread between those tables is visible *inside a single suite* β€” the cleanest
control in the project, because only one variable moves:

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/nikshepsvn/thimble/master/assets/catalog-control-dark.png">
  <img alt="Seal-Tools in-domain vs out-of-domain: row accuracy 33.1 vs 28.1, tool-name sequence 88.0 vs 79.0" src="https://raw.githubusercontent.com/nikshepsvn/thimble/master/assets/catalog-control-light.png" width="88%">
</picture>

Name-sequence accuracy tracks row accuracy exactly. The model is not failing to
extract arguments on unfamiliar catalogs β€” it is failing to pick the right
function.

**Disclosures.** Mobile Actions' public train split (8,693 rows, disjoint from
eval) is in the training mix β€” that is what the first table's heading means.
DroidCall's official split script calls `random.shuffle()` unseeded, so their
exact 200 rows are unreproducible by anyone; ours is a seeded split from the same
pool, firewalled out of training. The Seal-in margin over the yardstick is +0.5
on 700 rows, within sampling noise. The pre-registered dev-loss champion was a
sibling checkpoint scoring 28.4; that selector failure is diagnosed in
[FINDINGS.md](https://github.com/nikshepsvn/thimble/blob/master/FINDINGS.md) with both models' tables published.

## Adapting it to your catalog

```
python scripts/adapt.py --catalog my_tools.json --name mydomain
python scripts/eval_catalog.py --ckpt mydomain --baseline thimble-v6 \
    --catalog my_tools.json --gold my_eval.jsonl
```

Three stages, each resumable with `--stage`:

| stage | what happens |
|---|---|
| `synth` | a teacher model writes (query β†’ calls) rows **against your schemas**; each is validated against your parameter types and an evidence rule before it is kept |
| `pack` | your rows are blended with guard corpora and packed into two splits |
| `train` | continues from `thimble-v6`, annealing your blend into the LR-decay phase |

Two things about that recipe are load-bearing, both measured rather than assumed:

- **Anneal, don't retrain.** The same corrective corpus scored 28.4 fed from
  scratch and 33.1 annealed into the decay phase. Corrective data dilutes into
  the average when it competes with a whole corpus; it concentrates when it
  arrives late.
- **Keep the guard data.** The blend deliberately carries general tool-calling
  rows alongside yours. Annealing purely on your catalog trades away the
  competence you are building on. `adapt.py` warns if it finds none.

Pass `--examples` if you have real gold rows; they are weighted above synthetic
ones. Needs `OPENROUTER_API_KEY` for synthesis and a GPU to train. For scale, the
v6 cycle synthesized 74,250 validated rows for $56.

**Not yet demonstrated end to end.** `adapt.py` wires together exactly the
machinery that produced the v6 result, but no third-party catalog has been
adapted and published. The recipe is measured; the ergonomics are new. If you run
it, the numbers are worth a pull request.

## Deploying it

The Python stack is for training, eval and adaptation. To serve or embed the
model there is **[cengine/](https://github.com/nikshepsvn/thimble/tree/master/cengine)** β€” the full decoder (tokenizer, trunk,
grammar walk, name head, retrieval) in one dependency-free C file:

```
uv run python cengine/export.py && cd cengine && make
./thimble -w thimble-q8.bin -t tokenizer.bin -c demo_catalog.json \
    "make a reservation at Nobu for 2 people at 7pm"
```

| | weights | load | per request* |
|---|---:|---:|---:|
| Python stack (torch, CPU) | 184 MB | seconds | 582 ms |
| cengine fp32 | 191 MB | ~60 ms | 453 ms |
| cengine int8 | 48 MB | ~20 ms | 348 ms |

*mean over the first 100 Mobile Actions eval rows on an Apple M3; a request is
a full decode, prefill plus every choice point and rollback.

Parity is verified, not assumed: fp32 output is byte-identical to the Python
stack on 100/100 checked rows, and int8 differs on 2/100 β€” both of which
happened to move toward gold. Laptops, phones and edge Linux are in reach.

## Using it as an agent fast path

Semantic routers answer *which tool*; the request still pays an LLM call for
the arguments. **[route/](https://github.com/nikshepsvn/thimble/tree/master/route)** is the other half: a ~130ms local
dispatcher that returns the complete validated call plus two confidence
signals, so your big model is only consulted when the dispatcher abstains.

```python
from route.dispatch import ThimbleDispatcher

d = ThimbleDispatcher()   # wraps `cengine/thimble --serve`
r = d.dispatch("text Sam that i'm running late", tools)
r.dispatched   # True β€” confidence cleared the per-catalog gate
r.calls        # [{"name":"sendMessage","arguments":{"body":"i'm running late","contact":"Sam"}}]
```

The gate is measured, not assumed. On a catalog represented in training
(Mobile Actions, n=300 gold rows), sweeping the value-confidence threshold:

| gate | requests dispatched | precision of dispatched |
|---|---:|---:|
| vlp β‰₯ βˆ’0.002 | 77% | 98.7% |
| vlp β‰₯ βˆ’0.001 | 65% | 99.5% |

On a catalog the model handles poorly, the same gate collapses coverage to
~17% instead of dispatching confidently wrong calls. Derive the threshold for
*your* catalog from a small gold set (`route.dispatch.sweep`, one command);
if no threshold clears your bar, adapt the model first or keep everything on
the fallback path. A LangGraph node example with full router traceability is
in [route/langgraph_fastpath.py](https://github.com/nikshepsvn/thimble/blob/master/route/langgraph_fastpath.py).

## How it works

Most constrained-decoding systems bolt a grammar onto a model trained to generate
free text, then manage the mismatch. Here the **tokenizer, the training loss, and
the decoder are one design**, built around the same five decision points:
refuse-or-call, which tool, include this optional, what value, stop or continue.

**The tokenizer is built for the grammar.** JSON structural characters β€” and
digits β€” are singleton tokens. Structure can therefore be force-fed *exactly*,
with no token-healing and no ambiguity about where a constraint lands. The usual
arrangement masks logits over a vocabulary that merged `",` into a single token
and papers over the seam. Digits never merge either, so a copied number tokenizes
the same way every time; the rebuild was verified by a fragmentation gate
(word-value fragmentation 2.72 β†’ 2.34 tokens/word, digits lengthening by design).
It shipped as part of the v4 β†’ v5 bundle that took name-sequence accuracy from
80.4% to 91.5% β€” that bundle also added 350k corpus rows and reweighted the mix,
so the tokenizer's own share of the gain was never isolated.

**The loss is weighted by those same five decisions** β€” structure 1x, keys 1.5x,
names 2x, values 4x, stop-decision 6x β€” matched to the measured error
distribution. The model is optimized for the choices it will be asked to make,
not for tokens it will never emit. (A closely related weighting, without the
stop-decision term, appears independently in Needle's
[Simple Attention Networks notes](https://github.com/cactus-compute/needle/blob/main/docs/simple_attention_networks.md);
the scheme is not original to this project.)

**The decoder consults the model only at those points.** Everything else is
determined before it runs, which is where [the contract](#the-contract) comes
from. One call, start to finish β€” `MODEL` marks the only places the network is
asked anything:

```
  [                                    <- grammar
  └─ ? refuse or call ...................... MODEL
       β”‚
       β”œβ”€ refuse ──────────────► ]      <- grammar
       β”‚
       └─ call
          {"name":"                     <- grammar
          └─ ? which tool .................. MODEL
             ","arguments":{            <- grammar
             β”‚
             β”œβ”€ next key from YOUR schema  <- grammar
             β”‚  β”œβ”€ ? include it ........... MODEL
             β”‚  └─ ? what value ........... MODEL
             β”‚     (repeat for each key)
             β”‚
             }}                         <- grammar
             └─ ? stop or continue ........ MODEL
                β”œβ”€ continue ──► back to {"name":"
                └─ stop ──────► ]        <- grammar
```

Every `<- grammar` line is emitted without consulting the model at all. Argument
keys are iterated from your schema, which is why inventing one is not a
low-probability event β€” there is no step at which it could happen.

### Evidence the co-design works

Two measurements that look like caveats in isolation are the proof in context.

**There is no projection tax.** The same rows decoded with the grammar and with
no grammar at all (`scripts/draft_vs_constrained.py`):

| suite | free generation | grammar-constrained |
|---|---|---|
| Mobile Actions (150) | 78.7 | 78.7 |
| Seal-Tools in (150) | 26.7 | 28.0 |

On Mobile Actions the two agree on **150 of 150 rows**. The grammar is not
overriding the model β€” the model already wants what the grammar enforces. A
bolted-on grammar produces disagreement and a tax to recover; this is why
draft-then-constrain (DCCD) had nothing to recover here and was abandoned.

Stated plainly, because the distinction matters: the grammar buys *reliability*,
not accuracy. "Constrained decoding makes the model correct" would be a different
claim and not one this data supports. What it buys is that the worst failure
modes cannot be expressed, plus parseability on the ~11% of Seal rows where free
generation emits invalid JSON.

**And the co-design is load-bearing, not decorative.** Down-weighting the
grammar-forced tokens in the loss β€” on the theory that the model need not learn
what the decoder will supply β€” cost **12 points** in a controlled twin run. Those
tokens carry the call-sequencing signal: the model learns *when a call ends*
through structure it never has to emit. Remove them and it breaks.

<details>
<summary><b>The rest of the stack β€” retriever, name head, trunk</b></summary>


- **Retriever** β€” `retrieve(query, tools, emitted=...)`, a DTDR-style
  (arXiv 2512.17052) refresh conditioned on the *partial plan*, so the candidate
  set is recomputed after each emitted call rather than once per request.
- **Name head** β€” a bilinear readout scoring candidate tool-name spans in the
  prompt against the hidden state at the decision position. Selection is treated
  as pointing at the prompt, not generating from a vocabulary, following "Looking
  Is Not Picking" (arXiv 2606.16364): mis-selection is a readout failure, not a
  perception one. Its only positive result was on *unfamiliar* catalogs (+2.2),
  which is why it is on by default for your own tools.
- **Trunk** β€” deep-thin and gated: d=448, 20 layers, GQA 8/4, SwiGLU x2.0,
  QK-norm, sandwich RMSNorm, tied embeddings, Muon on 2D weights and AdamW on
  embeddings, norms and heads. This part is standard modern practice and is not
  where the advantage is; a controlled study from the Needle authors
  (arXiv 2607.18363) finds architecture choices at this scale worth hundredths of
  a nat at matched parameters. The co-design above is the part that matters.

</details>

<details>
<summary><b>How the model was built β€” the error-driven data loop</b></summary>


Row accuracy factors as `P(name sequence) x p^n`, where `p` is per-call argument
accuracy. Each version measured which factor was binding and attacked only that:

| version | name seq | p | Seal-in | what changed |
|---|---|---|---|---|
| v4 | 80.4% | 0.593 | 24.3 | baseline |
| v5 | 91.5% | 0.60 | 31.4 | 16k digit-singleton tokenizer, +350k corpus rows, seal_train x6, dev-selected EMA |
| **v6** | ~92% | ~0.66 | **33.1** | error-driven synth against three measured failure buckets, annealed into the decay phase |

The v6 data round came straight from the v5 diagnostic: of 193 failing calls, 66
added exactly one unmentioned optional, ~35 bound the wrong entity, ~30 missed
canonical date forms, ~29 were unwinnable noise in the gold. Mid-run causal
check: **+3.3 points at constant LR** from the corrective corpus alone. That loop
is what `adapt.py` automates for your catalog.

</details>

## What did not work

Eleven ideas, each killed by a measurement rather than an argument: span copying
(βˆ’30), pointer heads (βˆ’16), RFT-style loss down-weighting (βˆ’12), from-scratch
retraining (βˆ’4.7), field-set reranking (βˆ’1.4), beam/RL/best-of-N (oracle-capped
below target), draft-then-constrain (no tax to recover), a global optional-skip
prior (catalog-dependent), `MAX_CALLS` (never binding), RLOO on an annealed
checkpoint (diverges at every LR), and matching the benchmark's numeric typing
(not learnable). Plus two process failures that cost real points.

**[FINDINGS.md](https://github.com/nikshepsvn/thimble/blob/master/FINDINGS.md) has all of them** with the measurement, the reason,
and the takeaway. It is the most reusable part of the project.

## Known limits

- **Unfamiliar catalogs.** Out-of-domain name-sequence accuracy is 79% against
  88% in-domain. Every out-of-domain deficit traces to this number β€” the one
  `adapt.py` exists to move.
- **Optional arguments, in both directions.** The largest documented failure
  bucket: 66 of 193 failing v5 calls added exactly one optional the query never
  mentioned, and the model also drops optionals the query does state.
- **Multi-call tracks per-call accuracy, not call count.** `P(names) x p^n`, so
  chains collapse wherever `p` is mediocre and hold where it is not: 73.5% on
  Mobile Actions, 19.4% on Seal-Tools in-domain. The call count is not the
  problem; the catalog is.
- **Parallel calls are a separate, worse failure.** `parallel` 12.0,
  `live_parallel` 0.0 β€” repeated instantiations of one schema, as opposed to
  calls the query motivates in sequence.
- **Schema dialects.** `simple_python` 29.3 on BFCL against `simple_java` 14.0
  and `simple_javascript` 8.0. Those conventions are absent from a deliberately
  extractive ~1B-token corpus.
- **768-token context.** 151 of 3,641 BFCL rows (4.1%) do not fit and score as misses.
- **Microcontrollers.** ~11.5MB at 2-bit is a property of the parameter count,
  not a shippable artifact: 2-bit would need quantization-aware retraining this
  model never had. The smallest thing that actually runs is the 48MB int8
  engine ([Deploying it](#deploying-it)) β€” Pi-class and up, not Cortex-M.

Scale explains most of it honestly: ~1B unique tokens, no pretraining phase, a
corpus spent deliberately on depth rather than breadth.

## Going deeper

- **[FINDINGS.md](https://github.com/nikshepsvn/thimble/blob/master/FINDINGS.md)** β€” eleven negative results and two process
  failures, with the measurement that killed each one. The most reusable part.
- **[REPRODUCING.md](https://github.com/nikshepsvn/thimble/blob/master/REPRODUCING.md)** β€” repository layout and the exact
  pipeline that rebuilds the published numbers.
- **[RESULTS.md](https://github.com/nikshepsvn/thimble/blob/master/RESULTS.md)** β€” the full chronological experimental record.
- **[paper/thimble.pdf](https://github.com/nikshepsvn/thimble/blob/master/paper/thimble.pdf)** β€” the tech report: the co-design
  thesis, the negative results, and the related work, in citable form.

## Honest summary

Turning a request into calls against an API you control is a smaller problem than
the models usually pointed at it. Treated as a translation layer rather than a
language model, it fits in 48M parameters, comes with guarantees a prompted model
cannot offer, and can be specialized to one catalog for the price of a dinner.
It ships as a working system, not just a checkpoint: a one-file C engine with
byte-verified parity, a browser demo running the whole model client-side, and a
measured confidence gate for fronting a larger agent. Its limits are real and
measured rather than described. Built by one person over a few days with AI
assistance, for about the price of a video game console.

## Citation

```bibtex
@misc{saravanan2026thimble,
  title  = {Thimble: A 48M-Parameter Tool-Calling Model from Co-Designing the
            Tokenizer, Loss, and Decoder --- with Eleven Negative Results},
  author = {Saravanan, Nikshep},
  year   = {2026},
  url    = {https://github.com/nikshepsvn/thimble}
}
```