svara-tts-voiceclone-beta is an experimental extension of svara-tts-v1, designed to bring lightweight voice cloning and improved accent preservation to Indic languages. It introduces a simple but effective reference-swap finetuning technique, enabling more stable zero-shot speaker identity across long, expressive utterances.
Built on an Orpheus-style discrete audio token architecture, the model supports 19 languages, expressive cues (<laugh>, <yawn>, <angry>), and low-latency TTS on commodity hardware.
Research on speech identity, expressivity, and multilingual TTS
Out-of-Scope / Not Intended
Impersonating private individuals without consent
Fraud, targeted deception, harassment
High-risk or safety-critical deployments
Perfect 1:1 replication of voices (this is a beta research release)
Limitations
Zero-shot cloning is not identical to dedicated finetuning
Speaker similarity may degrade over long utterances
Varies by language due to dataset imbalance
Emotion emphasis may differ across low-resource languages
Rare names and numbers may require normalization or rewriting
These improve with targeted LoRA finetuning or higher-quality data.
Responsible Use
By using this model, you agree to follow applicable laws and ethical guidelines. Synthetic speech should be disclosed when appropriate. Avoid impersonation or harmful use cases.