mlboydaisuke commited on
Commit
d900b69
Β·
verified Β·
1 Parent(s): 6d0c68a

Add measured Galaxy S26 NPU/GPU section

Browse files

Numbers from the S26 classic sweep. Primary logs:
litertlm-convert/community_accel_work/s2_npu_sweep/results.jsonl (on-device
JIT) and portal_work/bench_all.txt + bench_redo.txt (the 2026-08-23 AOT
sweep); every row states which of the two produced it. Rows carry device,
SoC, runtime, N and thermal state, and join to this repo by exact .tflite
file name. Rows taken outside the sweep's thermal window are not reported.

Files changed (1) hide show
  1. README.md +17 -1
README.md CHANGED
@@ -27,7 +27,7 @@ tags:
27
  | File | Recipe | Size | Target |
28
  |---|---|---|---|
29
  | `LFM2.5-Encoder-350M-Prompt-Router_wi8fc.tflite` | int8 dynamic-range (linears + embedding, convs float) | 365 MB | mobile + desktop |
30
- | `LFM2.5-Encoder-350M-Prompt-Router_fp16.tflite` | fp16 weights, float compute | 713 MB | desktop β€” phone memory limits (XNNPACK per-signature fp32 unpacking) |
31
 
32
  Two signatures, `route_128` and `route_512` (S = 128 / 512, batch 1, right-padded, up to **8 lane slots**):
33
 
@@ -181,6 +181,22 @@ Android figures use the standard TFLite [`benchmark_model`](https://ai.google.de
181
 
182
  **GPU works as of the 2026-08-13 re-export.** The re-export respells the one idiom mobile GPU delegates refuse β€” transformers' rank-5 `repeat_kv` expand β€” into an equivalent rank-4 matmul (outputs bitwise-identical on CPU); the OpenCL delegate now takes the whole graph. Measured with the LiteRT CompiledModel API (fp32 GPU precision, real inputs incl. pooling matrices, best of 3 warm runs): `route_512` **18.7 ms** on the Pixel 8a, cosine **0.9949** vs the fp32 desktop reference; iPhone 17 Pro Metal `route_512` 172 ms, cosine **1.000000**. Set the GPU precision to **fp32** β€” at fp16 GPU precision this family's norm reductions overflow and every output is NaN. These CompiledModel timings are not comparable to the classic-delegate `benchmark_model` timings above (different GPU runtime).
183
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
184
  ## License
185
 
186
  LFM Open License v1.0 (see `LICENSE`, unchanged from the base model). Note the license's commercial-use threshold (Section 5). This repository redistributes converted **Derivative Works** of LiquidAI/LFM2.5-Encoder-350M-Prompt-Router with modification notices per Section 4; all credit for the model to [Liquid AI](https://www.liquid.ai/).
 
27
  | File | Recipe | Size | Target |
28
  |---|---|---|---|
29
  | `LFM2.5-Encoder-350M-Prompt-Router_wi8fc.tflite` | int8 dynamic-range (linears + embedding, convs float) | 365 MB | mobile + desktop |
30
+ | `LFM2.5-Encoder-350M-Prompt-Router_fp16.tflite` | fp16 weights, float compute | 713 MB | desktop CPU β€” phone CPU memory limits (XNNPACK per-signature fp32 unpacking); this is the file the Snapdragon NPU runs, AOT-compiled (see *Snapdragon NPU (Hexagon)*) |
31
 
32
  Two signatures, `route_128` and `route_512` (S = 128 / 512, batch 1, right-padded, up to **8 lane slots**):
33
 
 
181
 
182
  **GPU works as of the 2026-08-13 re-export.** The re-export respells the one idiom mobile GPU delegates refuse β€” transformers' rank-5 `repeat_kv` expand β€” into an equivalent rank-4 matmul (outputs bitwise-identical on CPU); the OpenCL delegate now takes the whole graph. Measured with the LiteRT CompiledModel API (fp32 GPU precision, real inputs incl. pooling matrices, best of 3 warm runs): `route_512` **18.7 ms** on the Pixel 8a, cosine **0.9949** vs the fp32 desktop reference; iPhone 17 Pro Metal `route_512` 172 ms, cosine **1.000000**. Set the GPU precision to **fp32** β€” at fp16 GPU precision this family's norm reductions overflow and every output is NaN. These CompiledModel timings are not comparable to the classic-delegate `benchmark_model` timings above (different GPU runtime).
183
 
184
+ ## Snapdragon NPU (Hexagon)
185
+
186
+ - `LFM2.5-Encoder-350M-Prompt-Router_fp16.tflite` β€” the NPU runs it at 176.1 ms. The GPU does not β€” `LiteRtException: Failed to compile model`.
187
+ - `LFM2.5-Encoder-350M-Prompt-Router_wi8fc.tflite` β€” the GPU runs it at 80.59 ms. The NPU does not β€” `LiteRtException: Failed to compile model`.
188
+
189
+ | file | backend | compiled | inference (median / min) | load |
190
+ |---|---|---|---:|---:|
191
+ | `LFM2.5-Encoder-350M-Prompt-Router_fp16.tflite` | NPU (Hexagon v81) | AOT (SM8850) | 176.1 ms / 171.8 ms | 434 ms |
192
+ | `LFM2.5-Encoder-350M-Prompt-Router_wi8fc.tflite` | GPU (Adreno) | β€” | 80.59 ms / 79.69 ms | 7676 ms |
193
+
194
+ Measured on a **Samsung Galaxy S26** (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT `CompiledModel` 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status `NONE` throughout. Headroom 0.75–0.78, where 1.0 is the throttling threshold.
195
+
196
+ The NPU row marked *AOT* ran an artifact compiled ahead of time for SM8850 (ai-edge-litert 2.2.0 + QAIRT 2.47.0), not the published file. That artifact is not distributed here; the compile is one command in the [NPU guide](https://github.com/john-rocky/hf-to-litertlm/blob/main/docs/android-npu.md).
197
+
198
+ GPU wiring: [GPU guide](https://github.com/john-rocky/hf-to-litertlm/blob/main/docs/android-gpu.md).
199
+
200
  ## License
201
 
202
  LFM Open License v1.0 (see `LICENSE`, unchanged from the base model). Note the license's commercial-use threshold (Section 5). This repository redistributes converted **Derivative Works** of LiquidAI/LFM2.5-Encoder-350M-Prompt-Router with modification notices per Section 4; all credit for the model to [Liquid AI](https://www.liquid.ai/).