jeffasante's picture
README: list f16, mark int8 as not published
75e735b verified
|
Raw
History Blame Contribute Delete
6.04 kB
metadata
license: gemma
library_name: cellm
pipeline_tag: text-generation
tags:
  - function-calling
  - tool-use
  - gemma3
  - quantization
  - on-device
  - cellm

functiongemma-270m β€” cellm builds

Function-calling Gemma3 (270M) converted to the .cellm format for the cellm CPU runtime.

File Size Recipe
functiongemma-270m-f16.cellm 511 MB unquantized reference
functiongemma-270m-int8e.cellm 257 MB int8 weights, int8 embedding
functiongemma-270m-int8-e4g32.cellm 186 MB int8 weights, int4 group-32 embedding
tokenizer.json 32 MB β€”

Architecture: gemma3_text, 18 layers, hidden 640, 4 query heads / 1 KV head, head_dim 256, vocab 262144, context 32768.

Which one to use

int8e (257 MB) unless you need the smaller file. It matches the f16 baseline on 11 of 16 prompts; int8-e4g32 matches on 9. The 71 MB saving costs two prompts and buys no speed.

Use int8-e4g32 (186 MB) only when the memory budget is binding.

Benchmark

16 prompts, greedy decoding (--temperature 0), 64 max new tokens, --stop-tokens 1,50,106. Apple Silicon, CPU backend. Two separate measures:

  • vs HF ref β€” exact generated-token-ID match against the original HuggingFace transformers model. Measures end-to-end correctness.
  • vs f16 β€” exact match against our own f16 .cellm build. Isolates quantization damage from pre-existing model behaviour.
Build Size vs HF ref vs f16 Avg prefill
f16 511 MB 12/16 16/16 4.19 s
int8 416 MB 11/16 12/16 2.95 s not published
int8e 257 MB 11/16 11/16 2.90 s recommended
int8-e4g32 186 MB 9/16 9/16 2.94 s

int8 is measured but not shipped: int8e matches it on every metric that matters β€” same 11/16, same speed β€” in 159 MB less.

Prefill is ~400 prompt tokens (the tool declarations dominate). Decode ran 12–32 tokens at roughly 26 tok/s.

Two honest caveats:

The f16 baseline is only 12/16. Four failures are inherited from the base model, not caused by quantization. Do not read "11/16" as "5 quantization regressions".

Quantization below int8 buys no speed. 257 MB and 186 MB are within noise of each other (2.90 s vs 2.94 s). The prefill path skips the lm_head matmul at every position except the last, so the embedding table β€” the only thing that differs between these two builds β€” is barely touched. Shrinking it saves memory, not time.

Sample outputs (int8-e4g32)

Correct:

Turn on wifi
  <start_function_call>call:set_wifi{enabled:true}<end_function_call>

What is the weather in Accra?
  <start_function_call>call:get_weather{city:<escape>Accra<escape>,unit:<escape>c<escape>}<end_function_call>

Turn off wifi and check the weather in London in fahrenheit
  <start_function_call>call:set_wifi{enabled:false}<end_function_call>
  <start_function_call>call:get_weather{city:<escape>London<escape>,unit:<escape>f<escape>}<end_function_call>

Who was the first president of Ghana?
  I apologize, but I cannot assist with historical inquiries about past
  leaders of countries. My current capabilities are focused on ...

Dropped calls β€” these are the two regressions unique to int8-e4g32; both emit only one of the two required calls:

Turn on wifi and set brightness to 50
  f16:        call:set_wifi{enabled:true} + call:set_brightness{level:50}
  int8-e4g32: call:set_wifi{enabled:true}                       <- brightness lost

Make the screen brighter, set it to 90
  f16:        call:set_brightness{level:90} + call:set_brightness{level:40}
  int8-e4g32: call:set_brightness{level:40}                     <- first call lost

Failures shared by every build, including f16 β€” base-model behaviour:

Wake me up at 6:30 tomorrow
  call:set_alarm{time:<escape>14:30<escape>}      <- wrong time, all builds

Text Ama that I'm running late
  call:send_message{...,recipient:<escape>person@example.com<escape>}
                                                  <- invents a recipient

Usage

infer \
  --model functiongemma-270m-int8e.cellm \
  --tokenizer tokenizer.json \
  --prompt "$PROMPT" \
  --gen 64 --temperature 0 \
  --stop-tokens 1,50,106

--stop-tokens 1,50,106 is required. The converter writes eos_token_id: 106 but the model actually stops on 50; without the explicit list, generation runs past the end of the function call.

The prompt must follow the Gemma chat template, with tool declarations wrapped in <start_function_declaration> / <end_function_declaration> before the <start_of_turn>user turn.

Quantization notes

These builds came out of a search for a sub-100 MB model. That target was not reached, and the reasons are worth recording:

The embedding tolerates 4 bits; the linear weights do not. With f16 weights, an int4 group-32 embedding scores 5/5 on the smoke set. With f16 embeddings, int4 group-32 weights score 0/5 β€” fluent but wrong text such as "I am sorry, but I cannot assist with this request." Hence the asymmetric recipe: int8 everywhere except the embedding.

int2 fails regardless of group size. The fixed codebook {-1.5, -0.5, 0.5, 1.5} holds relative reconstruction error at 0.43 on the embedding even at group 32 (0.471 per-row β†’ 0.431 at g32). Grouping cannot fix a codebook limit. Getting under 100 MB requires the 168M-parameter embedding at 2 bits, so the target is out of reach without a learned codebook.

The int4 scale is near-optimal, so that 10% weight error is inherent rather than a tuning artifact: amax/7 gives 0.1008 on mlp.up_proj at g32 versus 0.0967 for an exhaustively searched scale.

CPU only. Grouped scales are implemented in the CPU path. The Metal i4 path still passes hidden as the group size, so grouped models will produce wrong results on GPU. int8e is unaffected by this and works on both.

License

Gemma Terms of Use, inherited from the base model.