RTF - Real Time Factor - real time factor, how many seconds it takes to create 1 second of audio - (lower is
better)
WER - Word Error Rate - this is a proxy for understandability, assuming more understandable speech scores better
in STT, correlates to how many words the TTS pronounces wrong - (lower is better)
DAMERAU LEVENSHTEIN SIMILARITY - this is also a proxy for understandability, assuming more understandable speech… See the full description on the dataset page:
https://huggingface.co/datasets/Jarbas/ovos-tts-bench.