# BlueCodec edge-fixed encoder for the frozen official Supertonic-3 vocoder (Sept 2026) | File | Content | |------|---------| | `encoder.safetensors` | Encoder weights only (`encoder.*` keys), for `BlueCodec.from_pretrained(..., decoder="supertonic3", filename=...)` | | `LICENSE.OpenRAIL-M` | Copy of the license of the Supertonic-3 model this encoder was trained against | **The decoder is not in this repository.** Like `encoder_supertonic3_decoder/`, this encoder uses the official Supertonic-3 vocoder (`onnx/vocoder.onnx` from [Supertone/supertonic-3](https://huggingface.co/Supertone/supertonic-3), revision `3cadd1ee`, Supertone Inc., BigScience OpenRAIL-M). The `bluecodec` package downloads it from the official repo at load time: ```python from bluecodec import BlueCodec codec = BlueCodec.from_pretrained("notmax123/blue-codec", decoder="supertonic3", filename="encoder_supertonic3_decoder_edge_fixed/encoder.safetensors") # needs: pip install onnx z = codec.encode(audio, edge_pad_chunks=0) # official layout: pad to a multiple of 3072 samples, keep 6 * ceil(L / 3072) frames y = codec.decode(z)[..., :audio.shape[-1]] ``` **What changed.** The Supertonic-3-decoder encoder (`ae_290000.pt`) writes an "edge code" into the last latent frames of a clip that ends right after audio (last compressed frame ~9x the median norm). This encoder is that encoder continued for 60k steps (290k -> 350k) with the official vocoder still frozen, on edge-aware batches (variable-length segments, half ending at the clip's true end, batches padded to a multiple of 3072 samples, loss masked per clip at `ceil(L/3072) * 3072`) plus a tail-consistency term (relative L2 between each clip's last compressed latent frame and an encoding of the same clip with 2 extra chunks of silence). AdamW lr 2e-5 with a 1k-step warm-up and cosine decay; encoder, discriminators and optimizer moments continued from step 290k. Its latents are still in the Supertonic-3 latent space but are **not** identical to the 290k encoder's: stats files and downstream models built on one do not transfer to the other. **Audit** (10 reference clips, LibriTTS at 24 kHz resampled to 44.1 kHz; same metrics as the 290k encoder's card and the technical report): | Encoder + official vocoder | log-mel L1 (historical) | plain 228-band log-mel L1 | 12-22 kHz vs source | DNSMOS | |---|---:|---:|---:|---:| | Supertonic-3-decoder encoder (290k), bare encoding | 0.3385 | 0.8944 | +7.63 dB | 3.42 | | **Edge-fixed encoder (350k)**, official layout | **0.3387** | **0.8839** | **+7.11 dB** | **3.44** | End of clip, last compressed frame (normalised with each encoder's latent stats), 50 clips (the 10 references plus 40 training clips cut to native length, a multiple of 3072 samples, and multiples + 512 / + 1700): | Encoding | 290k: last / median norm | 290k: last frame vs edge-free (mean / max) | edge-fixed: last / median norm | edge-fixed: vs edge-free (mean / max) | |---|---:|---:|---:|---:| | official layout (`edge_pad_chunks=0`) | 1.76x | 99% / 363% | **1.36x** | **5.3% / 33%** (40/50 within 10%) | | edge-padded (`edge_pad_chunks=2`) | 1.36x | 2.3% / 4.2% | 1.37x | 0.5% / 0.9% | | edge-free reference (1 s silence appended) | 1.38x | 0 | 1.37x | 0 | With the official layout alone the end-of-clip spike is gone (last/median at the edge-free level). The remaining deviations come from clips whose audio runs to the very end of the array. `edge_pad_chunks=2` is still the closest to an edge-free encoding and is recommended when the latents feed another model. Details: https://github.com/maxmelichov/blue-codec (technical report, section 12). **License note.** This encoder contains none of Supertone's weights, but it was trained through their model, so it may be a "Derivative of the Model" under OpenRAIL-M ยง1(e). Its use is therefore subject to the use-based restrictions in Attachment A of `LICENSE.OpenRAIL-M` in addition to the MIT license of BlueCodec's own code and weights.