Dia
19.0kState-of-the-art dialogue TTS model. Generates ultra-realistic conversational speech with emotion, pauses, and multi-speaker support. 19k+ stars.
ai audio
Language: Python Org: @nari-labs Status: active
Dia is a 1.6B parameter text-to-dialogue model capable of generating highly realistic speech directly from text transcripts. Built at nari-labs, it supports multi-speaker generation, non-verbal cues like laughter and coughing, and emotion control — all from a single model. The architecture enables zero-shot voice cloning and fine-grained prosody control that outperforms commercial alternatives.