A unified speech-language model that synchronizes speech and text into a single, cohesive stream via 1:1 alignment.
Text-Acoustic Dual-Alignment Large Language Model
TADA is a unified speech-language model that synchronizes speech and text into a single, cohesive stream via 1:1 alignment. By leveraging a novel tokenizer and architectural design, TADA achieves high-fidelity synthesis and generation with a fraction of the computational overhead required by traditional models.