-
@coqui_ai @huggingface You say this is built on top of TorToiSe. I've done many months worth of experiments with it. Have gotten it to render in 4-5 seconds with decent quality, though, I have some questions.
-
@coqui_ai @huggingface When I tried the demo on huggingface, I noticed a non-negligible amount of unnatural stutters, and artifacting in the generated voice. The kind you see with a bad seed, and especially with a model trained on an incorrectly split dataset... Is this known?
-
@coqui_ai @huggingface What magic trickery is being used to get high quality out of the model here? What I've been doing is generating at a reliable and stable lower quality, then up-res'ing it after the fact.
-
@coqui_ai @huggingface I must assume the basis of the zero-shot voice cloning is still that of tortoise, since it shows similar limitations compared to fine-tune or post-generation style transfer...
-
@coqui_ai @huggingface Anyway, I was building out a TorToiSe-based high quality low latency TTS system of my own for a while, and maybe it might make more sense to provide input on this project instead. No need to spread out all the opensource when you can contribute, right?
AuroraNemoia’s Twitter Archive—№ 284