put the pth file in your SoVITS_weights_v2 folder and the ckpt in GPT_weights_v2
make both Language for Reference audio and Inference text language "Japanese", set slicing to "Slice by every punct".
the underlying model was trained on audio that was brickwall dynamic range compressed, like that commonly found in visual novels or video games.
you should be able to give it a CLEAN AND NOISE/MUSIC/STATIC free FEMALE Japanese voice clip from 3-10 seconds, give it a 100% ACCURATE transcription and get ok results out the other side.
I have found that results can be improved applying post-generation noise reduction and some treble boosting EQ processing. Audacity works well enough for this, since there isn't anything in the official Gradio interface.
Feel free to keep everything else at the defaults
If you want to start the inference engine auomatically, you can use do something like