IndexTTS-2.5 ONNX

This model is used by https://github.com/mercallureAI/local-multimodal-infra .

This repository provides an FP16 ONNX export of IndexTeam/IndexTTS-2.5.

Original model:

https://huggingface.co/IndexTeam/IndexTTS-2.5

Original project:

https://github.com/index-tts/index-tts

This is not an official IndexTTS release. It is intended for ONNX Runtime inference testing and deployment experiments.

Description

IndexTTS-2.5 is a zero-shot text-to-speech model that clones a voice from a single reference clip. It supports Chinese, English, Japanese, Spanish and Arabic, with emotion control and speaking-speed control.

The package was built from the original IndexTTS-2.5 PyTorch weights in two steps:

  1. Export: DakeQQ's exporter, used as is. The IndexTTS2 exporter (Index_TTS/v2) from DakeQQ/Text-to-Speech-TTS-ONNX was run without modification at commit f491c82. The graph split, the graph inputs and outputs, and the FP16 KV cache all come from this exporter. Our wrapper script made one change, and it is to transformers rather than to the exporter: it maps dtype= to torch_dtype= so the exporter runs on the transformers 4.52 that the official uv.lock pins.

  2. FP16 conversion and packaging: our own code. DakeQQ's Optimize_ONNX.py and its FP16 plan were not used. Our script:

    • converts the graphs with ONNX Runtime's convert_float_to_float16, keeping graph inputs and outputs in float32;
    • fuses the attention in the flow-matching DiT into com.microsoft.MultiHeadAttention;
    • zeroes a token-id sum in the Synthesis graph that overflows FP16;
    • removes redundant casts;
    • drops the graphs the runtime does not use.

    Weights shared between graphs are then deduplicated into IndexTTS_SharedInitializers.onnx.data (1,511 tensors, 2.78 GB) with DakeQQ's Shared_Weights.bundle_shared_initializers, used unmodified. Our script writes tokenizer.json and manifest.json.

The graphs cover reference-audio processing, the GPT backbone, the flow-matching mel decoder and the BigVGAN vocoder. Text normalization, tokenization, segmentation of long text and the sampling loop run outside the graphs, in your own code.

Files

File Precision Role
IndexTTS2_ReferencePreprocess.onnx Mixed: w2v-BERT encoder FP16, audio front end and speaker/mel path FP32 Reference audio (22.05 kHz mono) to w2v-BERT semantic features, CAMPPlus speaker style and the reference mel condition
IndexTTS2_Conditioning.onnx FP16 Speaker latent and emotion vector, from the speaker and emotion reference features, the 8-way emotion weights and emotion_alpha
IndexTTS2_TargetPrefill_sampling.onnx FP16 GPT prefill over the text tokens and language_id; returns the FP16 KV cache and the first sampled mel code
IndexTTS2_DecodeStep_sampling.onnx FP16 One autoregressive step with top-k, top-p, temperature and repetition penalty; repeat until the stop code 8193
IndexTTS2_Synthesis.onnx FP16 Mel codes to the flow-matching condition; applies duration_factor and cfg_rate
IndexTTS2_CFMEstimator.onnx FP16 One flow-matching step; run 25 times (cfm_steps)
IndexTTS2_Decoder.onnx FP16 BigVGAN vocoder: mel to 22.05 kHz waveform
IndexTTS2_Metadata.onnx Carries the package contract as ONNX metadata
IndexTTS_SharedInitializers.onnx + .onnx.data FP16, some FP32 Weight store referenced by the graphs above
tokenizer.json Tokenizer description: split pattern, special tokens and language ids
multilingual_zh_ja_yue_char_del.tiktoken Text vocabulary, copied from the original model
manifest.json Package contract, runtime constants, shared-tensor layout and SHA-256 of every file
LICENSE, LICENSE_ZH bilibili Model Use License Agreement (English and Chinese)
config.json Package summary; the Hub counts downloads by requests for this file. Not read at run time

Keep all files in one directory: the graphs load their weights from IndexTTS_SharedInitializers.onnx.data by relative path.

Runtime notes

  • Requires ONNX Runtime: IndexTTS2_CFMEstimator.onnx uses the com.microsoft.MultiHeadAttention contrib operator. The graphs use opset 20.
  • Graph inputs and outputs are float32, except the KV cache, which is float16.
  • Disable the ONNX Runtime optimizers CastFloat16Transformer and FuseFp16InitializerToFp32NodeTransformer when creating sessions (listed in manifest.json under disabled_optimizers).
  • FP16 targets NVIDIA GPUs (CUDA execution provider); plan for about 6 GB of VRAM. CPU works but is slow.
  • Audio input and output are 22,050 Hz. Text is at most 600 tokens per segment; split longer text and join the segments.
  • Emotion is an 8-float vector in the order [happy, angry, sad, afraid, disgusted, melancholic, surprised, calm], scaled by emotion_alpha. Emotion from a text description is not included: the QwenEmotion graphs were not exported.
  • Only the sampling variants of the prefill and decode graphs are included; greedy decoding is not.
  • Pronunciation can be fixed inline as <word|READING>, using Pinyin with tone numbers, CMU phonemes or Kana, as in the original project.

Source

Item Revision
IndexTeam/IndexTTS-2.5 weights c39ce5ba981572cb187443877ff559dfb246ce63
index-tts/index-tts code ee40fa7d6c6b8a2c7f06105f9f1e65775b74868c
DakeQQ/Text-to-Speech-TTS-ONNX exporter (unmodified) f491c82ee6b29dd760c3c00e25552be8925358a7

DakeQQ/Text-to-Speech-TTS-ONNX is licensed under Apache-2.0. Only its output is distributed here; this repository contains none of its code.

The graphs also contain weights of the auxiliary models that IndexTTS-2.5 loads at runtime: facebook/w2v-bert-2.0, funasr/campplus and nvidia/bigvgan_v2_22khz_80band_256x. Their own licenses also apply.

Testing

The package was run end to end with ONNX Runtime from Rust on Chinese, English, Japanese and Spanish text, and a Chinese sample was transcribed back with ASR as a smoke check. Arabic has not been tested. No listening test or benchmark against the PyTorch model has been run, and FP16 output is not bit-identical to the original.

License

This repository is derived from IndexTeam/IndexTTS-2.5 and is distributed under the original bilibili Model Use License Agreement. See LICENSE and LICENSE_ZH.

The license is not Apache-2.0. Among other terms, organizations above the user or revenue thresholds in section 2.2 must request a separate license from bilibili, a copy of the agreement must be kept with every copy of the model, and downstream recipients must comply with it.

All rights to the original model, code and training work belong to the original authors.

Disclaimer

This is an unofficial ONNX export of IndexTTS-2.5.

It is not affiliated with, maintained by, or endorsed by the original IndexTTS authors.

Do not use this model to clone a person's voice without their consent, or for fraud, impersonation or misinformation.

Citation

@misc{li2026indextts25technicalreport,
      title={IndexTTS 2.5 Technical Report},
      author={Yunpei Li and Xun Zhou and Jinchao Wang and Lu Wang and Yong Wu and Siyi Zhou and Yiquan Zhou and Yining Wang and Yaogen Yang and Zhetao Hu and Shiyao Duan and Jiacheng Xu and Bin Xia and Jingchen Shu},
      year={2026},
      eprint={2601.03888},
      archivePrefix={arXiv},
      primaryClass={cs.SD},
      url={https://arxiv.org/abs/2601.03888},
}
Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ModaLeap/indextts-2.5-onnx

Quantized
(3)
this model

Paper for ModaLeap/indextts-2.5-onnx