MioCodec's speaker encoder, in one file for the browser
The global-embedding path of Aratako/MioCodec-25Hz-24kHz's encoder β reference audio in, a 128-dimensional speaker embedding out β collected into a single f32 safetensors so a browser can fetch it in one request. Built for voxshot.
This is a repackaging, not a new model. No weight here was trained; every tensor is copied from one of the two sources below.
encoder-weights.safetensors |
117,303,232 B, 99 tensors, 29,322,838 params |
encoder-weights.json |
8,905 B β tensor index and provenance |
Why it exists
The weights for this one path live in two different places: the
GlobalEncoder (50 tensors) is in MioCodec's own checkpoint, while the WavLM
front end it runs on top of is a torchaudio download. A browser should not
have to know that, nor fetch a 378 MB .pth to use 4 MB of it. So the path is
collected once, ahead of time, into one file.
Only the global branch is included. MioCodec's encoder also produces content tokens β that half is not here, and is not needed for speaker embedding.
What the path is
24 kHz mono
β symmetric zero pad
β resample 24kβ16k (torchaudio's polyphase filter, precomputed
here as a conv1d kernel [2, 1, 23], stride 3)
β WavLM feature extractor (7 strided convs, GroupNorm on the first, GELU)
β LayerNorm + Linear 512β768
β positional conv (grouped k=128, weight-norm already folded)
β transformer layers 1 and 2 (12 heads; bucketed, gated relative-position
bias β the global branch reads only these two)
β mean of the two layers
β ConvNeXt backbone, 4 blocks, 768β384
β attentive statistics pooling β Linear 768β128 β LayerNorm
β 128-dim embedding
Two details a reimplementation gets wrong and this file cannot tell you, so they are written down here:
- torchaudio's Base+ WavLM is post-norm, so an encoder-level LayerNorm runs between the positional-conv residual and layer 1.
- the pooling softmax is over time, not channels.
Naming
WavLM tensors keep torchaudio's own names under a wavlm. prefix, the
GlobalEncoder tensors keep MioCodec's names under global_encoder., and the
precomputed resampler is resample.kernel. encoder-weights.json lists every
tensor with its shape.
The positional convolution's weight normalization is already folded into
its weight β there is no separate weight_g / weight_v to combine.
Provenance
Produced by spike/miocodec/export_encoder_weights.py from:
Aratako/MioCodec-25Hz-24kHz model.safetensors
523,087,956 bytes
sha256 60483759cde136451d53ff5a2f7e283c015758ab540de9037b5d4519366a7705
torchaudio WAVLM_BASE_PLUS wavlm_base_plus.pth
377,604,347 bytes
sha256 136a3e720c04f2c77bf7a4dc6a3868b14d5a2c145a988114b733cb1a8428be98
Both digests are recorded in encoder-weights.json and cross-checked at export
time against the golden this port is verified with, so the weights and the
checkpoints they came from cannot drift apart unnoticed.
Accuracy
The TypeScript port that consumes this file was checked stage by stage against the reference encoder running under torch, on two real reference clips and one synthetic input. Every stage agrees to 1e-4 relative to its own peak, except four ConvNeXt stages on one clip which reach 1e-4 in the reference's own f32-versus-f64 comparison β conditioning of that clip, not port error. The final embedding agrees to 7.3e-7.
License and attribution
MIT, which both sources carry.
- MioCodec: Aratako/MioCodec-25Hz-24kHz by Chihiro Arata, MIT.
- WavLM Base+: originally published by the authors of WavLM (Chen et al., 2022) under the MIT License and redistributed by torchaudio under the same license. Source Β· License
@article{chen2022wavlm,
title = {WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing},
author = {Chen, Sanyuan and Wang, Chengyi and Chen, Zhengyang and others},
journal = {IEEE Journal of Selected Topics in Signal Processing},
year = {2022}
}
Model tree for m96-chan/MioCodec-encoder-webgpu
Base model
Aratako/MioCodec-25Hz-24kHz