TeamFolklore_Model / README.md
SherTheCoder's picture
Upload 13 files
3dc8e68 verified
|
Raw History Blame Contribute Delete
22.7 kB
---
license: apache-2.0
language:
- pa
- hi
library_name: tflite
pipeline_tag: audio-classification
tags:
- keyword-spotting
- wake-word
- streaming
- depthwise-separable-cnn
- tinyml
- microcontroller
- edge-ai
- int8
- stm32
- esp32
- punjabi
- hindi
datasets:
- SherTheCoder/TeamFolklore_UthoKikkar
---
<p align="center">
<img src="assets/banner.png" alt="The kikkar flower, from bud to full bloom" width="100%">
</p>
<h1 align="center">Kikkar KWS</h1>
<p align="center">
<b>A 44 KB int8 streaming model that listens for "Utho Kikkar" on a microcontroller, and for nothing else.</b><br>
We made it for projects and prototypes with less than 256 KB of free RAM
</p>
<p align="center">
<img src="https://img.shields.io/badge/task-keyword%20spotting-2a78d6" alt="Task: keyword spotting">
<img src="https://img.shields.io/badge/format-TFLite%20int8-4a3aa7" alt="Format: TFLite int8">
<img src="https://img.shields.io/badge/size-44%20KB-1baf7a" alt="Size: 44 KB">
<img src="https://img.shields.io/badge/runs%20on-STM32N657%20%7C%20ESP32-e8a33d" alt="Runs on STM32N657 and ESP32">
</p>
<table align="center">
<tr>
<td align="center" width="25%"><h2>44 KB</h2>in flash</td>
<td align="center" width="25%"><h2>32,849</h2>parameters</td>
<td align="center" width="25%"><h2>45,872</h2>multiply-accumulates per step</td>
<td align="center" width="25%"><h2>5</h2>layers</td>
</tr>
<tr>
<td align="center"><h2>30 ms</h2>per step</td>
<td align="center"><h2>≈ 1.5 s</h2>of context</td>
<td align="center"><h2>≈ 38 KB</h2>working memory</td>
<td align="center"><h2>60 ms</h2>phrase end to wake-up</td>
</tr>
</table>
## Overview & Wake-Word Detection
Our model listens in steps of 30 ms. At every step it takes the three newest
frames of log-mel energy, updates what it remembers of the last one and a half
seconds, and gives back one number: the chance that someone has just finished
saying *Utho Kikkar* (ਉੱਠੋ ਕਿੱਕਰ · उठो किक्कर), which means "wake up, Kikkar". We run it
entirely on the board, and no audio leaves the device until it fires.
We trained it to fire only at the end of the whole phrase. We taught it that
*Utho* on its own, *Kikkar* on its own, the two the wrong way round, and
everyday words that sound close, like *chakkar*, *shakkar* or *kicker*, are all
things that are not the wake word.
It is the listener inside Kikkar, our Smart India Hackathon 2026 project. When
it fires, a flower blooms on the board's screen and we stream whatever you say
next to a speech recogniser in the cloud.
| At a glance | |
|---|---|
| **Input** | 3 new log-mel frames × 40 bands per 30 ms step, int8 |
| **Output** | one probability per step: has "Utho Kikkar" just ended? |
| **Architecture** | Streaming depthwise-separable CNN, 5 layers |
| **Size** | 32,849 parameters, ≈ 44 KB of flash |
| **Compute** | 45,872 multiply-accumulates per step, 1.53 million a second |
| **Context** | ≈ 1.5 s of audio |
| **Latency** | 60 ms from the end of the phrase to the wake-up |
| **Format** | TensorFlow Lite, int8, streaming state kept inside the model |
| **Runs on** | STM32N657 (Cortex-M55) and ESP32 |
| **Trained on** | About 2,000 real recordings plus synthetic speech, including [SherTheCoder/TeamFolklore_UthoKikkar](https://huggingface.co/datasets/SherTheCoder/TeamFolklore_UthoKikkar) |
## Quickstart & Verification Pipeline
```python
# pip install numpy soundfile tensorflow huggingface_hub
import numpy as np
import soundfile as sf
import tensorflow as tf
from huggingface_hub import hf_hub_download
path = hf_hub_download("SherTheCoder/TeamFolklore_Model", "kikkar_kws_int8.tflite")
interp = tf.lite.Interpreter(model_path=path)
interp.allocate_tensors()
inp, out = interp.get_input_details()[0], interp.get_output_details()[0]
```
We wrote the front end below to follow our model's specification. Features that are even
slightly different will quietly cost accuracy, so if your own pipeline differs
in any detail, trust that instead.
```python
SR, WIN, HOP, NFFT, MELS, STEP = 16000, 480, 160, 512, 40, 3
def _mel(hz):
return 2595.0 * np.log10(1.0 + hz / 700.0)
def _hz(mel):
return 700.0 * (10.0 ** (mel / 2595.0) - 1.0)
_edges = _hz(np.linspace(_mel(20.0), _mel(7600.0), MELS + 2))
_bins = np.arange(NFFT // 2 + 1) * SR / NFFT
_fbank = np.zeros((MELS, NFFT // 2 + 1))
for b in range(MELS):
lo, mid, hi = _edges[b:b + 3]
rise = (_bins > lo) & (_bins < mid)
fall = (_bins >= mid) & (_bins < hi)
_fbank[b, rise] = (_bins[rise] - lo) / (mid - lo)
_fbank[b, fall] = (hi - _bins[fall]) / (hi - mid)
_hann = np.hanning(WIN)
def features(audio, cms_seconds=3.0):
"""16 kHz mono int16 in, 40 normalised log-mel values per 10 ms frame out."""
x = np.asarray(audio, dtype=np.float64) / 32768.0
frames = np.array([x[s:s + WIN] * _hann for s in range(0, len(x) - WIN + 1, HOP)])
power = np.abs(np.fft.rfft(frames, NFFT)) ** 2
logmel = np.log(np.maximum(power @ _fbank.T, 1e-8)) # floor at -80 dB
a = 1.0 - np.exp(-(HOP / SR) / cms_seconds)
mean = logmel[:50].mean(axis=0) # start from the first 0.5 s
out = np.empty_like(logmel)
for i, frame in enumerate(logmel):
mean += a * (frame - mean)
out[i] = frame - mean
return out
def wake_scores(audio):
"""One probability per 30 ms step, fed to the model the way the board feeds it."""
feats = features(audio)
interp.reset_all_variables() # fresh streaming state
scale, zero = inp["quantization"]
scores = []
for s in range(0, len(feats) - STEP + 1, STEP):
x = feats[s:s + STEP]
if inp["dtype"] == np.int8:
x = np.clip(np.round(x / scale) + zero, -128, 127)
interp.set_tensor(inp["index"], x.astype(inp["dtype"]).reshape(inp["shape"]))
interp.invoke()
scores.append(float(interp.get_tensor(out["index"]).reshape(-1)[0]))
return np.array(scores)
```
Our model ends in a float sigmoid, so each score is already a probability. Then
point it at a clip:
```python
audio, sr = sf.read("clip.wav", dtype="int16")
assert sr == SR and audio.ndim == 1, "the model expects 16 kHz mono audio"
scores = wake_scores(audio)
smoothed = np.convolve(scores, np.ones(3) / 3)[:len(scores)] # the last 3 steps, 90 ms
print("highest smoothed score:", round(float(smoothed.max()), 3))
```
To see where the board would wake up (on the board we use a threshold of 0.7,
picked on the int8 model):
```python
def detect(audio, threshold, refractory_s=1.0):
"""Times, in seconds, at which the board would wake up."""
scores = wake_scores(audio)
smoothed = np.convolve(scores, np.ones(3) / 3)[:len(scores)]
wakes, last = [], -np.inf
for i, p in enumerate(smoothed):
t = ((i * STEP + STEP - 1) * HOP + WIN) / SR
if p >= threshold and t - last >= refractory_s:
wakes.append(round(t, 2))
last = t
return wakes
```
## Audio Frontend & Feature Extraction
<p align="center">
<img src="assets/features.png" alt="16 kHz audio, 30 ms Hann windows every 10 ms, 512-point power spectrum, 40 mel bands from 20 Hz to 7.6 kHz, log and running-mean normalisation, then 3 frames per 30 ms model step" width="100%">
</p>
| Stage | Setting | Output |
|---|---|---|
| Capture | 16 kHz, mono, 16-bit | 160 samples every 10 ms |
| Window | 30 ms Hann (480 samples), every 10 ms | 480 samples |
| FFT | 512-point, power spectrum | 257 bins |
| Mel | 40 triangular bands, 20 Hz to 7.6 kHz | 40 values |
| Compression | natural log, floored at −80 dB | 40 values |
| Normalisation | per-band running mean subtracted, τ ≈ 3 s | 40 × int8 |
| Model step | the 3 newest frames, every 30 ms | 3 × 40 |
We rely on the running mean to soak up the differences between microphones,
gain settings and rooms, so we do not need to recalibrate the board for every
house. We start it from the first half second of audio after boot.
## Model Architecture & Layer Specifications
We built a streaming depthwise-separable CNN in five layers. The first two work across
frequency while that axis still exists, learning spectral shapes cheaply. The
third folds frequency into a single 48-channel vector. Everything after that is
causal depthwise convolution over time plus 1×1 pointwise mixing, which we chose
because it streams as a small ring-buffer update at every step and is the int8
path that microcontroller kernels run fastest.
<p align="center">
<img src="assets/architecture.png" alt="Input of 3 new log-mel frames by 40 bands; S1 frequency stem to 3 by 20 by 16; S2 depthwise-separable block to 1 by 10 by 32; S3 collapse from 320 values to 48 channels; six causal temporal blocks at 48 channels; head from 48 to 1 with a sigmoid" width="100%">
</p>
### Layer Breakdown & Compute (MACs)
Shapes are per 30 ms step. We make every convolution over time causal, padded on
the left only, and every convolution over frequency keeps its size.
| # | Layer | Operation | Output (T, F, C) | Weights | MACs per step |
|---|---|---|---|---:|---:|
| S1 | Frequency stem | Conv2D 3×3, stride (1, 2), 1 → 16, BN, ReLU | 3, 20, 16 | 144 | 8,640 |
| S2a | Frequency depthwise | Depthwise Conv2D 3×3, stride (3, 2), 16 channels, BN, ReLU | 1, 10, 16 | 144 | 1,440 |
| S2b | Channel mix | Pointwise 16 → 32, BN, ReLU | 1, 10, 32 | 512 | 5,120 |
| S3 | Collapse | Flatten 10 × 32 = 320, 1×1 conv to 48, BN, ReLU | 1, 48 | 15,360 | 15,360 |
| B1 | Temporal block, d = 1 | Depthwise Conv1D k5, pointwise 48 → 48, residual | 1, 48 | 2,544 | 2,544 |
| B2 | Temporal block, d = 1 | same | 1, 48 | 2,544 | 2,544 |
| B3 | Temporal block, d = 2 | same | 1, 48 | 2,544 | 2,544 |
| B4 | Temporal block, d = 2 | same | 1, 48 | 2,544 | 2,544 |
| B5 | Temporal block, d = 3 | same | 1, 48 | 2,544 | 2,544 |
| B6 | Temporal block, d = 3 | same | 1, 48 | 2,544 | 2,544 |
| H | Head | Pointwise 48 → 1, sigmoid in float | 1, 1 | 48 | 48 |
<p align="center">
<img src="assets/temporal-block.png" alt="Inside each temporal block: a causal depthwise 1D convolution with kernel 5 and dilation d, batch norm and ReLU, a pointwise 48 to 48 convolution and batch norm, added back to the input, then ReLU" width="100%">
</p>
We picked dilations of 1, 1, 2, 2, 3, 3, which give the trunk 49 steps of
context, 1.47 s, and the stem and the window overhang add about 40 ms more:
≈ 1.51 s in all. We sized that to cover *Utho Kikkar*, which takes about 1.0 to
1.2 s to say, with room for slow speakers, and no more. Extra context would cost RAM and let the model latch on
to background sound.
<p align="center">
<img src="assets/compute.png" alt="Multiply-accumulates per step: S3 collapse 15,360 (33.5 percent), temporal blocks 15,264 (33.3 percent), S1 stem 8,640 (18.8 percent), S2 block 6,560 (14.3 percent), head 48 (0.1 percent)" width="100%">
</p>
| Parameters | Count |
|---|---:|
| Convolution weights, S1 to S3 | 16,160 |
| Convolution weights, B1 to B6 | 15,264 |
| Head weights | 48 |
| **Convolution subtotal** | **31,472** |
| Batch-norm scale and bias, folded (688 channels × 2) | 1,376 |
| Head bias | 1 |
| **Total** | **32,849** |
At 33.3 steps a second our model does 1.53 million multiply-accumulates a
second.
## RAM Allocation & Tensor Arena Budget
<p align="center">
<img src="assets/memory.png" alt="Design budget for the listening path: pre-roll ring 32 KB, TFLM tensor arena about 20 KB, task stacks 8 KB, FFT scratch 4 KB, interpreter about 3 KB, I2S DMA 1.3 KB, frame buffer 1.0 KB, mel and CMS state 0.4 KB, smoothing 0.1 KB" width="100%">
</p>
| Component | RAM |
|---|---:|
| I2S DMA buffers, 4 × 10 ms × int16 | 1.3 KB |
| Frame assembly buffer, 480 samples | 1.0 KB |
| FFT scratch, 512 complex float32 | 4.0 KB |
| Mel output and CMS running state | 0.4 KB |
| TFLM tensor arena: activations, kernel scratch and 2.3 KB of streaming state | ≈ 20 KB |
| Interpreter, op resolver and allocator | ≈ 3 KB |
| Smoothing and refractory state | 0.1 KB |
| Task stacks, audio and inference | 8 KB |
| **Keyword spotting in total** | **≈ 38 KB** |
| Pre-roll ring, 1 s at 16 kHz int16, for the hand-off to the ASR server | 32 KB |
| **Listening path in total** | **≈ 70 KB** |
That leaves us about 186 KB of the 256 KB limit for everything else. The peak
activation is only 960 bytes, at S1's output; the arena is mostly kernel scratch
and alignment.
### Ring Buffers & Streaming State
We keep this state between steps, inside the TFLM arena as resource variables,
so we allocate nothing separately.
| Buffer | Shape | Bytes |
|---|---|---:|
| S1 time history | 2 frames × 40 | 80 |
| S2a time history | stride 3 over kernel 3, nothing to keep | 0 |
| B1 ring, d = 1 | 4 × 48 | 192 |
| B2 ring, d = 1 | 4 × 48 | 192 |
| B3 ring, d = 2 | 8 × 48 | 384 |
| B4 ring, d = 2 | 8 × 48 | 384 |
| B5 ring, d = 3 | 12 × 48 | 576 |
| B6 ring, d = 3 | 12 × 48 | 576 |
| **Total** | | **2,384** |
A dilated causal depthwise convolution needs (k − 1) × d × C bytes of history
at int8.
### Flash
| Item | Bytes |
|---|---:|
| int8 weights | 31,472 |
| int32 biases, 689 channels × 4 | 2,756 |
| Per-channel scales, 689 × 4 | 2,756 |
| FlatBuffer structure and metadata | ≈ 4,096 |
| Mel filterbank and Hann window tables | ≈ 3,072 |
| **Total** | **≈ 44 KB** |
## Post-Processing & Trigger Logic
<p align="center">
<img src="assets/trigger.png" alt="Illustration: the smoothed score stays low while Utho Kikkar is said, rises once it ends, and crosses the threshold 60 ms later, which wakes the board; a second peak inside the 1 s refractory period is ignored" width="100%">
</p>
<p align="center"><sub>An illustration of the rule, not a recording.</sub></p>
| Step | Setting |
|---|---|
| Model output | one probability per 30 ms step |
| Smoothing | average of the last 3 steps, 90 ms |
| Threshold τ | 0.7, picked from the DET curve of the int8 model |
| Refractory | 1 s after a wake-up |
| Phrase end to wake-up | 60 ms: up to 30 ms for the step that holds the end of the phrase, then one more step |
The wake-up comes quickly because of how we labelled the model. We turn its
target on one step *before* the phrase ends, so by the time the step holding the
end is done, two of the three steps in the average are already high, and one
more step takes it over τ. Computing a step takes a millisecond or less, so it
barely adds to the 60 ms.
Quantisation moves the operating point, so we always pick the threshold on the
int8 model and never carry it over from the float one. If false activations
will not come down, the cheapest remaining lever we have is two or three frames
of lookahead, paid for directly in latency.
## Training Methodology & Supervision
<p align="center">
<img src="assets/supervision.png" alt="One label per 30 ms step: for Utho Kikkar only the six steps around the end of the phrase are labelled 1, and mid-phrase steps are 0; wrong order and half a phrase are 0 throughout" width="100%">
</p>
**Phrase-completion supervision.** We built the head as a per-step detector,
not a clip classifier. We label only the six steps from one before the end of
the phrase to four after it, about 180 ms, as 1. We label every other step 0,
including the middle of the phrase, and partial or reversed phrases 0
throughout. That is what teaches the model word order, and it makes the firing
delay equal to the smoothing window. We get the phrase end from energy-based
silence trimming on recordings, or straight from the synthesiser for TTS clips,
and when we speed a clip up or slow it down, we move its end by the same factor.
| Frames | Label |
|---|---|
| Steps from t_end − 1 to t_end + 4, about 180 ms | 1 |
| Every other step, mid-phrase included | 0 |
| Partial phrase, *utho* or *kikkar* alone | 0 throughout |
| Reversed order, *kikkar utho* | 0 throughout |
**Loss and selection.** We use per-step weighted binary cross-entropy.
Positives are about 6 steps in 300, but a positive weight of 50 over-fires, so
we start it at 8; we weight confusable negatives 8 and easy ones 1. We choose
checkpoints on the fewest false activations per hour of held-out continuous
speech first, and on recall only among those that meet the false-activation
target, never on accuracy.
**Positives.** We use about 2,000 real recordings, split by speaker so that
the model never hears validation and test speakers in training. On top of that,
we add several thousand synthetic renditions from AI4Bharat Indic-TTS, IndicF5,
Indic Parler-TTS, MMS-TTS in Punjabi and Hindi, and Piper, with varied prosody
and gaps between the two words. We oversample the real recordings so they carry
25 to 40% of the positive signal, otherwise the model learns vocoder artefacts.
**Confusable negatives.**
| Kind | Examples |
|---|---|
| Partial phrase | *utho*, *kikkar* |
| Wrong order | *kikkar utho* |
| Sounds like *kikkar* | kukkad, kukkar, chakkar, takkar, shakkar, nikkar, fikar, Makkar, Thakkar, kirkiri |
| Sounds like *utho* | utha, uthao, uthe, utha lo, uncha |
| Natural continuations | utho ji, utho beta, utho jaldi |
| Across word boundaries | "…jhoota kikar…", "…peeche kikar…" |
| English in the middle | kicker, quicker, sticker, liquor, auto, photo |
**Background negatives.** We use continuous Hindi and Punjabi speech from
Common Voice, IndicVoices and Shrutilipi, plus sounds from around the house, and
stream them as long segments so that training sees what the board hears.
**Augmentation.** We apply room impulse responses (OpenSLR-28, BUT ReverbDB) at
different distances; noise from MUSAN, DEMAND and ESC-50 at 0 to 20 dB SNR; gain from −20
to 0 dBFS; speed 0.9 to 1.1; pitch ±2 semitones; microphone response tilt; mild
clipping; frequency masking. We keep time masking light, because the labels
depend on timing.
**Hard negatives.** We train, stream the model over more than 10 hours of
Hindi, Punjabi and English speech, collect every false trigger as a new
negative, and train again, two or three times over.
**Streaming.** We train with causal convolutions over long 8 to 10 s chunks,
which is fast and parallel, then deploy the same weights running step by step
on ring buffers. We check that the streaming and full-chunk outputs agree
within 1e-3 before quantising.
**Quantisation.** We quantise after training, with a representative set of
real streaming features rather than synthetic clips, and fold batch norm into
the convolutions before export.
| Tensor | Scheme |
|---|---|
| Activations | int8, asymmetric, per tensor |
| Weights | int8, symmetric, per channel |
| Bias | int32 |
| Final sigmoid | float32, on a single number |
## Hardware Profiling & Benchmarks
<table>
<tr>
<td align="center" width="50%"><img src="assets/board-stm32n657.jpg" alt="STM32N6570-DK with the kikkar flower in full bloom on its display" width="100%"></td>
<td align="center" width="50%"><img src="assets/board-esp32.jpg" alt="ESP32 board with an I2S microphone on perfboard" width="100%"></td>
</tr>
<tr>
<td align="center"><b>Premium:</b> STM32N657, Ethernet and an 800 × 480 display</td>
<td align="center"><b>Budget:</b> ESP32 and an I2S microphone, on Wi-Fi</td>
</tr>
</table>
<p align="center">
<img src="assets/on-device.png" alt="Idle CPU: STM32N657 0.5%, ESP32 4%, limit 10%. RAM used: STM32N657 150 KB, ESP32 190 KB, limit 256 KB. Handoff latency: STM32N657 80 ms, ESP32 90 ms." width="100%">
</p>
| | STM32N657, premium | ESP32, budget | SIH limit |
|---|---|---|---|
| CPU while idling in continuous listening | **0.5%** | **4%** | under 10% |
| RAM used | **150 KB** | **190 KB** | under 256 KB |
| Latency, keyword end to audio at the ASR server | **80 ms** | **90 ms** | as low as possible |
These are our numbers for the whole application, networking and display
included, not the model alone. Of the 80 ms, 60 ms is our model making sure; the other
20 ms is the board getting the first audio to the server over Ethernet. Wi-Fi
takes 30 ms for the same step on the ESP32.
### Core Utilization & RTOS Tasks
On an ESP32-S3 at 240 MHz with ESP-NN int8 kernels:
| Task | Rate | Share of one core |
|---|---|---|
| FFT, mel and CMS | 100 a second | ≈ 0.6% |
| Inference | 33.3 a second | 0.7 to 2.6% |
| I2S, DMA and housekeeping | | ≈ 0.5% |
| **Idle listening in total** | | **≈ 2 to 4%** |
We run capture and inference as separate FreeRTOS tasks pinned to different cores.
A classic ESP32 has no vector unit and runs the inference about three times
slower, which still fits well inside 10% of a core.
## Target Firmware Integration (STM32 & ESP32)
On the ESP32 we run the model under TensorFlow Lite for Microcontrollers, with
int8 kernels from ESP-NN. Its 2.4 KB of streaming state lives inside the TFLM
arena as resource variables, so we have nothing extra to allocate. Size the
arena from `arena_used_bytes()` measured on the device, not from the estimate
above.
On the STM32N657 we do not use TFLM at all. We wrote a small exporter that turns
this `.tflite` into a flat list of 23 operations, and our firmware runs them with
its own int8 kernels, using Helium vector instructions for the dot products. It keeps the
same 2,384 bytes of history in rings, one per layer that looks back in time,
and the whole model works inside a 5.3 KB arena. Our firmware, the exporter,
the flower and the streaming pipeline are on [GitHub](https://github.com/SherTheCoder/Folklore).
## Field Testing & Edge Constraints
- **Test it on real people in real rooms,** and pick the threshold there, on
the int8 model.
- **It knows one phrase.** Ours is not a general speech model, and our model cannot
currently tell who is speaking.
- **The front end has to match.** The running-mean normalisation is part of
the input, and it needs its first half second after boot to settle.
- **Privacy is the point of our design.** We keep all audio on the board
until the wake word is confirmed.
## Licence
We release it under the [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
licence.
## Citation
```bibtex
@misc{teamfolklore2026kikkarkws,
title = {Kikkar KWS: a 44 KB streaming wake-word model for Utho Kikkar},
author = {{Team Folklore}},
year = {2026},
howpublished = {\url{https://huggingface.co/SherTheCoder/TeamFolklore_Model}}
}
```
## Made by
We are Team Folklore, and we made this for Smart India Hackathon 2026. Our
firmware, flower and streaming pipeline are on [GitHub](https://github.com/SherTheCoder/Folklore);
our training data is [SherTheCoder/TeamFolklore_UthoKikkar](https://huggingface.co/datasets/SherTheCoder/TeamFolklore_UthoKikkar).