IndexTTS 2.5 Technical Report

Index SpeechTeam

Abstract

In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-based Text-to-Semantic (T2S) module and a non-autoregressive Semantic-to-Mel (S2M) module, which together enable faithful emotion replication and establish the first autoregressive duration-controllable generative paradigm. Building upon this, we present IndexTTS 2.5, which significantly enhances multilingual coverage, inference speed, and overall synthesis quality through four key improvements:

Experiments show that IndexTTS 2.5 not only supports broader language coverage but also replicates emotional prosody in unseen languages under the same zero-shot setting. IndexTTS 2.5 achieves a 2.28× improvement in real-time factor (RTF) while maintaining comparable word error rate (WER) and speaker similarity to IndexTTS 2.

IndexTTS2.5: Controllable Emotional Speech Generation for Audiovisual Dubbing – A Case Study on Iconic Scenes from The Mermaid
IndexTTS2.5: Controllable Emotional Speech Generation for Audiovisual Dubbing – A Case Study on Iconic Scenes from Infernal Affairs
IndexTTS2.5: Controllable Emotional Speech Generation for Audiovisual Dubbing – A Case Study on Iconic Scenes from Pegasus

Contents

1 Zero-shot In-context Generation

IndexTTS 2.5 supports zero-shot voice cloning in Chinese, English, Japanese, Spanish and Arabic, preserving speaker identity from a single reference audio. All samples use neutral prosody.

Language Audio-Prompt Text Model Audio
ZH
同形观音座莲为观音座莲科观音座莲属下的一个种。 IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
如果你愿意,请想象一个曼彻斯特人、一个波士顿人、一个牙买加人和一个悉尼人围坐在多伦多的一家餐馆里吃饭。 IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
EN
Animal Liberation and the Royal Society for the Prevention of Cruelty to Animals (RSPCA) are again calling for the mandatory installation of CCTV cameras in all Australian abattoirs. IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
The U.S. says it has received information from an undisclosed source that specifically mentions the use of suicide bombers to will blow up "prominent landmarks" in Ethiopia and Kenya. IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
JA
このマインドセットの焦点は、スピード、論理性、正確性にあり、事実の特定、既存のテクニックの再適用、情報収集にも力を入れています。 IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
あるとき、恐怖に怯える王妃の前で、暴徒の1人がヴェルサイユ宮殿で殺された王室の衛兵の首を振って見せた。 IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
ES
Tras la derrota, Blanc anunció su dimisión. IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
Por lo general, siempre que se mueva dentro de esta región podrá atravesar las fronteras sin pasar otra vez por un puesto de control de pasaportes. IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
AR
بل أنتم في خواتيمه في أفضل لياليه وأيامه. هذه عشره الأخيرة. IndexTTS 2.5
CosyVoice 3 N/A
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
متسائلة ماذا لو كنا رواية تحكى؟ كيف ستكون بدايتها وماذا عن نهايتها؟ تفاصيل كثيرة تجول بيننا وبين ذواتنا. تفاصيل البعض منها جميل. ونتمنى لو أننا نملك حق إعادة تلك الأوقات. IndexTTS 2.5
CosyVoice 3 N/A
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2

2 Emotionally Expressive Speech Generation

IndexTTS 2.5 accurately replicates the emotional prosody present in the prompt audio across all supported languages. The following examples demonstrate emotion reproduction in Chinese, English, and the newly supported Japanese, Spanish, and Arabic. Note that due to the lack of Arabic emotional audio, we used Japanese audio as prompt here.

Language Emotion Audio-Prompt Text Audio
ZH Angry
我说,你别叨叨了行不行,你这辈子还没叨叨够哇?
Sad
每次闭上眼睛,脑海里全是那些让人痛心的回忆,我忍不住流泪。
Happy
祝你十四岁快乐! 愿您的生活永远充满阳光和欢笑!
Neutral
镰叶山龙眼为山龙眼科山龙眼属下的一个种。
EN Happy
I was chatting with an old classmate just now, reminiscing about the silly things we did as kids. We laughed so hard, tears almost came out. I really miss those days!
Sad
I feel like I’m lost in the darkness and can’t find a way out anymore.
Angry
If you’d ever treated me like a human being, would they have dared to do this?
Neutral
The tower caused minor discontent, because it blocked sight lines of Central Park.
JA Happy
ちょうど探しに行こうかなって思っていたんだ。どうかな、一緒に練習でも。
Neutral
小さい頃、異邦から流れ着いた勇者が剣術修行に付き合ってくれたり、魔物を倒しながら一緒に世界を救う旅をしたり…なんてのをよく想像していたんだ。
Guilty
いつ咲こうとも、風に吹き落される時は…来る。
Sad
いつ咲こうとも、風に吹き落される時は…来る。
ES Happy
No esperaba que esto fuese tan útil. Ahora que has obtenido un ascenso. ¿Puedes mejorar mi habitación? Gracias por valorar alguien tan vulnerable como yo, Doc. Seguiré dando lo mejor de mí.
Excited
¿Te has perdido? No te preocupes. Dime adónde quieres ir. Para llegar allí, sube a las escaleras de la izquierda del cuarto piso. Luego, gira a la derecha, y traspasa la tercera cruce, vea a la izquierda. Baja el piso, y sigue atravesar a la izquierda. Está justo enfrente de la sala de ingeniería de número tres.
Surprise
No es tu dieta de hacer los bajos de tu ropa y tus zapatos están limpios. ¡Muy bien! Espera, llevas el cuello de la ropa muy arrugado. No pretenderás ir a la fiesta así, ¿no? Vuelve y arregla eso. Te espero aquí.
Neutral
Ehh, ¿me das un momento para asimilarlo? Supongo que ya es tarde para echarse atrás. Vamos allá.
AR Happy
أنا في قمة السعادة اليوم، كل شيء يسير على ما يرام!
Angry
هذا الموقف سخيف ومستفز لأقصى درجة.
Sad
الدموع لا تفارق عيني، أشتاق له كثيراً.
Neutral
هذه هي النسخة النهائية من الوثيقة.

3 Cross-lingual Generation

IndexTTS 2.5 supports generating speech using prompt audio from another language (we only demonstrate those using Chinese prompt audio here), while maintaining the accent and prosody of the target language.

Language Audio-Prompt Text Model Audio
EN
Elmore is also a lawyer, having received a law degree from Northwestern University. IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
JA
傷の上を、手ぬぐいで冷やされると、ずいぶんしみたけれども、周作は我慢をした。 IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
ES
Los especialistas detectaron la formación de cristales en el orín de los felinos, por la incorporación de melamina y ácido cianúrico. IndexTTS 2.5
CosyVoice 3
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2
AR
والوفاء بالعهود والصبر والإنفاق ومقابلة السيئة بالحسنة وفيها التحذير من بعض انفصال الموجبة للطعن والطرد من رحمة الله تعالى IndexTTS 2.5
CosyVoice 3 N/A
Fish Audio S2 Pro
MossTTS v1.5
OmniVoice
VoxCPM2

4 Speaker and Emotion Disentangle

IndexTTS 2.5 decouples speaker identity from emotional expression, allowing users to independently specify a timbre prompt and an emotion prompt. The timbre prompt determines the speaker's voice characteristics, while the emotion prompt controls the emotional style of the generated speech. The two prompts can come from different speakers or even different languages.

Language Timbre-Prompt Emotion-Prompt Text Audio
ZH
尾号四四九幺的乘客刚夸了你,厉害了我的师傅,你真是个活地图。
EN
Have you visited that famous bat cave, Albert?
JA
電気について知らないのは非常識とされるようになる。
ES
Esto se consigue eliminando información redundante.
AR
فكلُّها خياراتٌ اتَّخذتَها بنفسِك، كشخصٍ بالغٍ يتحمّلُ مسؤوليّتَه كاملةً.

5 Emotion Vector Control

IndexTTS 2.5 supports controlling emotion intensity through an emotion weight parameter, allowing fine-grained control over the expressiveness of generated speech.

Language Audio-Prompt Emotion Category Emotion Weight Text Audio
ZH
Calm 0.5 有些人走了就再也没有回来过,所以等待和犹豫是这个世界上最无情的杀手。
0.8
1.0
Angry 0.5 我站在人海中,却感觉比任何时候都要孤独。
0.8
1.0
EN
Disgusted 0.5 Have you visited that famous bat cave, Albert?
0.8
1.0
Happy 0.5 He decided to spend the night there.
0.8
1.0
JA
Happy 0.5 翻訳ツールのおかげで外国のサイトもよく見るようになった。
0.8
1.0
Angry 0.5 翻訳ツールのおかげで外国のサイトもよく見るようになった。
0.8
1.0
ES
Melancholic 0.5 En la actualidad se continúa utilizando con mucha frecuencia una corona abierta.
0.8
1.0
Angry 0.5 Los esfuerzos policiales para encontrarla no arrojaron ningún resultado.
0.8
1.0
AR
Happy 0.5 فكلُّها خياراتٌ اتَّخذتَها بنفسِك، كشخصٍ بالغٍ يتحمّلُ مسؤوليّتَه كاملةً
0.8
1.0
Melancholic 0.5 الطقس جميل اليوم، هيا نخرج ونستمتع!
0.8
1.0

6 Grapheme-to-Phoneme (G2P) Control

IndexTTS 2.5 provides fine-grained pronunciation control via language-specific Grapheme-to-Phoneme (G2P) interfaces: Pinyin for Chinese, the CMU Pronouncing Dictionary for English, and Kana for Japanese. This enables precise disambiguation of polyphones and accurate rendering of context-dependent pronunciations.

Language Prompt Text With G2P Without G2P
ZH
认为反孔并非<掊|POU3>击孔子思想本身,乃<掊|POU3>击专制政治之灵魂
ZH
只见会计<长|ZHANG3>正着急地数着单据:“这次出差差不多花了万把块,我得好好核对!“
EN
He had a <minute|M IH1 . N AH0 T> to examine the <minute|M AY0 . N UW1 T> details of the contract.
EN
<Bilibili|B IY1 . L IY1 . B IY1 . L IY1> is one of China's most popular video platforms.
JA
彼は料理が<上手|じょうず>だが、囲碁では<上手|うわて>に負けた。