Upload only a voice you own or have permission to clone. A clean,
single-speaker 3–12 second clip works best.
Consent warning: do not impersonate people or clone a voice without
the speaker's explicit authorization.
🚧 Public preview: the interface is published, but generation will be enabled only after the checkpoint passes native-speaker listening and held-out evaluation.
0.81.2
Examples
Exact reference transcript
Haryanvi text
Speed
Model adaptation is experimental and does not represent every Haryanvi dialect.