notmax123's picture
Add encoder_supertonic3_decoder_edge_fixed/ (edge-fixed encoder for the official Supertonic-3 vocoder; encoder only)
ad9991c verified
|
Raw History Blame Contribute Delete
4.01 kB

BlueCodec edge-fixed encoder for the frozen official Supertonic-3 vocoder (Sept 2026)

File Content
encoder.safetensors Encoder weights only (encoder.* keys), for BlueCodec.from_pretrained(..., decoder="supertonic3", filename=...)
LICENSE.OpenRAIL-M Copy of the license of the Supertonic-3 model this encoder was trained against

The decoder is not in this repository. Like encoder_supertonic3_decoder/, this encoder uses the official Supertonic-3 vocoder (onnx/vocoder.onnx from Supertone/supertonic-3, revision 3cadd1ee, Supertone Inc., BigScience OpenRAIL-M). The bluecodec package downloads it from the official repo at load time:

from bluecodec import BlueCodec
codec = BlueCodec.from_pretrained("notmax123/blue-codec", decoder="supertonic3",
                                  filename="encoder_supertonic3_decoder_edge_fixed/encoder.safetensors")   # needs: pip install onnx
z = codec.encode(audio, edge_pad_chunks=0)   # official layout: pad to a multiple of 3072 samples, keep 6 * ceil(L / 3072) frames
y = codec.decode(z)[..., :audio.shape[-1]]

What changed. The Supertonic-3-decoder encoder (ae_290000.pt) writes an "edge code" into the last latent frames of a clip that ends right after audio (last compressed frame ~9x the median norm). This encoder is that encoder continued for 60k steps (290k -> 350k) with the official vocoder still frozen, on edge-aware batches (variable-length segments, half ending at the clip's true end, batches padded to a multiple of 3072 samples, loss masked per clip at ceil(L/3072) * 3072) plus a tail-consistency term (relative L2 between each clip's last compressed latent frame and an encoding of the same clip with 2 extra chunks of silence). AdamW lr 2e-5 with a 1k-step warm-up and cosine decay; encoder, discriminators and optimizer moments continued from step 290k. Its latents are still in the Supertonic-3 latent space but are not identical to the 290k encoder's: stats files and downstream models built on one do not transfer to the other.

Audit (10 reference clips, LibriTTS at 24 kHz resampled to 44.1 kHz; same metrics as the 290k encoder's card and the technical report):

Encoder + official vocoder log-mel L1 (historical) plain 228-band log-mel L1 12-22 kHz vs source DNSMOS
Supertonic-3-decoder encoder (290k), bare encoding 0.3385 0.8944 +7.63 dB 3.42
Edge-fixed encoder (350k), official layout 0.3387 0.8839 +7.11 dB 3.44

End of clip, last compressed frame (normalised with each encoder's latent stats), 50 clips (the 10 references plus 40 training clips cut to native length, a multiple of 3072 samples, and multiples + 512 / + 1700):

Encoding 290k: last / median norm 290k: last frame vs edge-free (mean / max) edge-fixed: last / median norm edge-fixed: vs edge-free (mean / max)
official layout (edge_pad_chunks=0) 1.76x 99% / 363% 1.36x 5.3% / 33% (40/50 within 10%)
edge-padded (edge_pad_chunks=2) 1.36x 2.3% / 4.2% 1.37x 0.5% / 0.9%
edge-free reference (1 s silence appended) 1.38x 0 1.37x 0

With the official layout alone the end-of-clip spike is gone (last/median at the edge-free level). The remaining deviations come from clips whose audio runs to the very end of the array. edge_pad_chunks=2 is still the closest to an edge-free encoding and is recommended when the latents feed another model.

Details: https://github.com/maxmelichov/blue-codec (technical report, section 12).

License note. This encoder contains none of Supertone's weights, but it was trained through their model, so it may be a "Derivative of the Model" under OpenRAIL-M §1(e). Its use is therefore subject to the use-based restrictions in Attachment A of LICENSE.OpenRAIL-M in addition to the MIT license of BlueCodec's own code and weights.