|
Download encoder_supertonic3_decoder_edge_fixed/README.md from notmax123/blue-codec: direct link, hf CLI and curl.
- Browser
- Download file 4.01 kB
-
https://huggingface.co/notmax123/blue-codec/resolve/main/encoder_supertonic3_decoder_edge_fixed/README.md
- Command line
-
hf download hf://notmax123/blue-codec/encoder_supertonic3_decoder_edge_fixed/README.md
-
curl -L -o README.md https://huggingface.co/notmax123/blue-codec/resolve/main/encoder_supertonic3_decoder_edge_fixed/README.md
4.01 kB
| # BlueCodec edge-fixed encoder for the frozen official Supertonic-3 vocoder (Sept 2026) | |
| | File | Content | | |
| |------|---------| | |
| | `encoder.safetensors` | Encoder weights only (`encoder.*` keys), for `BlueCodec.from_pretrained(..., decoder="supertonic3", filename=...)` | | |
| | `LICENSE.OpenRAIL-M` | Copy of the license of the Supertonic-3 model this encoder was trained against | | |
| **The decoder is not in this repository.** Like `encoder_supertonic3_decoder/`, this encoder uses the official Supertonic-3 vocoder (`onnx/vocoder.onnx` from [Supertone/supertonic-3](https://huggingface.co/Supertone/supertonic-3), revision `3cadd1ee`, Supertone Inc., BigScience OpenRAIL-M). The `bluecodec` package downloads it from the official repo at load time: | |
| ```python | |
| from bluecodec import BlueCodec | |
| codec = BlueCodec.from_pretrained("notmax123/blue-codec", decoder="supertonic3", | |
| filename="encoder_supertonic3_decoder_edge_fixed/encoder.safetensors") # needs: pip install onnx | |
| z = codec.encode(audio, edge_pad_chunks=0) # official layout: pad to a multiple of 3072 samples, keep 6 * ceil(L / 3072) frames | |
| y = codec.decode(z)[..., :audio.shape[-1]] | |
| ``` | |
| **What changed.** The Supertonic-3-decoder encoder (`ae_290000.pt`) writes an "edge code" into the last latent frames of a clip that ends right after audio (last compressed frame ~9x the median norm). This encoder is that encoder continued for 60k steps (290k -> 350k) with the official vocoder still frozen, on edge-aware batches (variable-length segments, half ending at the clip's true end, batches padded to a multiple of 3072 samples, loss masked per clip at `ceil(L/3072) * 3072`) plus a tail-consistency term (relative L2 between each clip's last compressed latent frame and an encoding of the same clip with 2 extra chunks of silence). AdamW lr 2e-5 with a 1k-step warm-up and cosine decay; encoder, discriminators and optimizer moments continued from step 290k. Its latents are still in the Supertonic-3 latent space but are **not** identical to the 290k encoder's: stats files and downstream models built on one do not transfer to the other. | |
| **Audit** (10 reference clips, LibriTTS at 24 kHz resampled to 44.1 kHz; same metrics as the 290k encoder's card and the technical report): | |
| | Encoder + official vocoder | log-mel L1 (historical) | plain 228-band log-mel L1 | 12-22 kHz vs source | DNSMOS | | |
| |---|---:|---:|---:|---:| | |
| | Supertonic-3-decoder encoder (290k), bare encoding | 0.3385 | 0.8944 | +7.63 dB | 3.42 | | |
| | **Edge-fixed encoder (350k)**, official layout | **0.3387** | **0.8839** | **+7.11 dB** | **3.44** | | |
| End of clip, last compressed frame (normalised with each encoder's latent stats), 50 clips (the 10 references plus 40 training clips cut to native length, a multiple of 3072 samples, and multiples + 512 / + 1700): | |
| | Encoding | 290k: last / median norm | 290k: last frame vs edge-free (mean / max) | edge-fixed: last / median norm | edge-fixed: vs edge-free (mean / max) | | |
| |---|---:|---:|---:|---:| | |
| | official layout (`edge_pad_chunks=0`) | 1.76x | 99% / 363% | **1.36x** | **5.3% / 33%** (40/50 within 10%) | | |
| | edge-padded (`edge_pad_chunks=2`) | 1.36x | 2.3% / 4.2% | 1.37x | 0.5% / 0.9% | | |
| | edge-free reference (1 s silence appended) | 1.38x | 0 | 1.37x | 0 | | |
| With the official layout alone the end-of-clip spike is gone (last/median at the edge-free level). The remaining deviations come from clips whose audio runs to the very end of the array. `edge_pad_chunks=2` is still the closest to an edge-free encoding and is recommended when the latents feed another model. | |
| Details: https://github.com/maxmelichov/blue-codec (technical report, section 12). | |
| **License note.** This encoder contains none of Supertone's weights, but it was trained through their model, so it may be a "Derivative of the Model" under OpenRAIL-M §1(e). Its use is therefore subject to the use-based restrictions in Attachment A of `LICENSE.OpenRAIL-M` in addition to the MIT license of BlueCodec's own code and weights. | |