Apps & Tools
IndexTTS-2.5
A zero-shot multilingual text-to-speech model with emotion and speed control.
Playable
IndexTTS-2.5 media is blocked
Allow external media to connect to the provider and play this content.
Description
IndexTTS-2.5 is a zero-shot voice cloning model supporting Chinese, English, Japanese, Spanish, and Arabic. It features an autoregressive GPT backbone with emotion control disentangled from timbre and controllable speaking speed, requiring roughly 6 GB of VRAM for inference.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.