Commit History

tflite: bake output limiter into SAME-S/SAME-L decoders (ceiling 0.977, matches TRT graft)
da6edc5
verified

cortexelus commited on

tflite: bake output limiter into SAME-S/SAME-L decoders (ceiling 0.977, matches TRT graft)
7d85335
verified

cortexelus commited on

model card: note optimized TFLite/LiteRT rung codecs
7f26210
verified

cortexelus commited on

tflite: SAME codecs -> static rung models (fp32 + w8a8); retire w16a32/w8a32/w8a8-dyn to legacy/
73fa327
verified

cortexelus commited on

tflite: preserve pre-rung SAME codecs under legacy/
416a9d0
verified

cortexelus commited on

Move the retired fp8_fast ONNX under onnx/same-s/legacy/
b5182df
verified

cortexelus commited on

Rename the autoencoder ONNX to match the engine naming scheme
4ea87d6
verified

cortexelus commited on

Move the retired w8_bf16 ONNX under onnx/*/legacy/
74cd2b7
verified

cortexelus commited on

Delete the miscalibrated sm_120 SAME-S fp8 encoder
197c684
verified

cortexelus commited on

Move deprecated sm_120 TRT engines under legacy/
19ab22a
verified

cortexelus commited on

Move deprecated sm_90 TRT engines under legacy/
ceaecd4
verified

cortexelus commited on

build-ready ONNX: limiter grafted + ceiling baked (decoders), amax-recalibrated activation scales (fp8 encoders)
d070aac
verified

cortexelus commited on

build-ready ONNX: limiter grafted + ceiling baked (decoders), amax-recalibrated activation scales (fp8 encoders)
b1400fd
verified

cortexelus commited on

build-ready ONNX: limiter grafted + ceiling baked (decoders), amax-recalibrated activation scales (fp8 encoders)
3f689ec
verified

cortexelus commited on

build-ready ONNX: limiter grafted + ceiling baked (decoders), amax-recalibrated activation scales (fp8 encoders)
f29bd0e
verified

cortexelus commited on

build-ready ONNX: limiter grafted + ceiling baked (decoders), amax-recalibrated activation scales (fp8 encoders)
9aa2165
verified

cortexelus commited on

build-ready ONNX: limiter grafted + ceiling baked (decoders), amax-recalibrated activation scales (fp8 encoders)
dcaaf84
verified

cortexelus commited on

retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders)
3963f0e
verified

cortexelus commited on

retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders)
f061405
verified

cortexelus commited on

retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders)
4b4cf7c
verified

cortexelus commited on

retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders)
935685c
verified

cortexelus commited on

retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders)
276f110
verified

cortexelus commited on

retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders)
0996c8a
verified

cortexelus commited on

retire fp8_fast; remove the superseded sm_90 fp8 engines (miscalibrated encoder, 8192-ceiling decoders)
cf76999
verified

cortexelus commited on

SAME-L/SAME-S: bake the limiter ceiling so the chunkable decoders keep the ('latent','pcm') signature; add the SAME-S fp8 chunkable pair
17258a7
verified

cortexelus commited on

SAME-L/SAME-S: bake the limiter ceiling so the chunkable decoders keep the ('latent','pcm') signature; add the SAME-S fp8 chunkable pair
7ca4180
verified

cortexelus commited on

SAME-L/SAME-S: bake the limiter ceiling so the chunkable decoders keep the ('latent','pcm') signature; add the SAME-S fp8 chunkable pair
775a981
verified

cortexelus commited on

SAME-L/SAME-S: bake the limiter ceiling so the chunkable decoders keep the ('latent','pcm') signature; add the SAME-S fp8 chunkable pair
aea1cf0
verified

cortexelus commited on

SAME-L/SAME-S: bake the limiter ceiling so the chunkable decoders keep the ('latent','pcm') signature; add the SAME-S fp8 chunkable pair
9a01fee
verified

cortexelus commited on

SAME-S: add the chunkable bf16 pair (two profiles, limiter)
baa3cd7
verified

cortexelus commited on

SAME-S: add the chunkable bf16 pair (two profiles, limiter)
7c73bf8
verified

cortexelus commited on

SAME-L encoders: low band 64 -> 256 (1.3-1.5x faster, cos 0.99831 -> 0.99969)
1a8702f
verified

cortexelus commited on

SAME-L encoders: low band 64 -> 256 (1.3-1.5x faster, cos 0.99831 -> 0.99969)
d0d486b
verified

cortexelus commited on

SAME-L: remove miscalibrated enc_fp8.trt (activation scales 1.3-28x too small; superseded by enc_fp8_chunkable.trt on sm_90)
2c8b02a
verified

cortexelus commited on

SAME-L: remove miscalibrated enc_fp8.trt (activation scales 1.3-28x too small; superseded by enc_fp8_chunkable.trt on sm_90)
e00128c
verified

cortexelus commited on

SAME-L: add enc_fp8_chunkable.trt (fp16, two-profile chunkable)
7036ca2
verified

cortexelus commited on

SAME-L: add dec_fp8_chunkable_limiter.trt (fp16, two-profile chunkable, runtime limiter ceiling)
093c4d6
verified

cortexelus commited on

SAME-L: add enc_fp16_chunkable.trt (fp16, two-profile chunkable)
eb343c9
verified

cortexelus commited on

SAME-L: add dec_fp16_chunkable_limiter.trt (fp16, two-profile chunkable, runtime limiter ceiling)
5dba282
verified

cortexelus commited on

cpu-amx README: document the bf16 DiT tier (--dit-precision)
6736003
verified

cortexelus commited on

cpu-amx: add dit_medium_bf16_core.bin (bf16 DiT engine)
0905190
verified

cortexelus commited on

cpu-amx: add dit_medium_bf16_flash.so (bf16 DiT engine)
60952d8
verified

cortexelus commited on

cpu-amx: add dit_medium_bf16_pin_fp32.npz (bf16 DiT engine)
476830e
verified

cortexelus commited on

cpu-amx: add dit_medium_bf16_core_manifest.txt (bf16 DiT engine)
70aafed
verified

cortexelus commited on

cpu-amx: add dit_medium_bf16.so (bf16 DiT engine)
fc19859
verified

cortexelus commited on

Delete old fp16mixed names + retired medium dit_bf16.trt (fp16 rename shipped)
48e34e2
verified

cortexelus commited on

Rename fp16mixed->fp16 DiT/T5Gemma engines+onnx (copies; deletes follow after code merge)
3f29675
verified

cortexelus commited on

Remove redundant w8_bf16 engines (int8 weights fold to bf16 at build → identical size/speed to the bf16 baseline)
63cc70d
verified

cortexelus commited on

Add sm_120 TRT engines (AOT SWA, wide profile) for the quantized tiers
4e8cc21
verified

cortexelus commited on