Views
No views yet
1KPipeline benchmark for voice af_heart (warm-up took 0.175s) using hexgrad/kokoro
2Test Chars Output (s) Inf(s) RTFx Peak GB
31 42 2.750 0.187 14.737x 1.44
42 129 8.625 0.530 16.264x 1.85
53 254 15.525 0.923 16.814x 2.65
64 93 6.125 0.349 17.566x 2.66
75 104 7.200 0.410 17.567x 2.70
86 130 9.300 0.504 18.443x 2.72
97 197 12.850 0.726 17.711x 2.83
108 6 1.350 0.098 13.823x 2.83
119 1228 76.200 4.342 17.551x 3.19
1210 567 35.200 2.069 17.014x 4.85
1311 4615 286.525 17.041 16.814x 4.78
14Total - 461.650 27.177 16.987x 4.85 PYTORCH_ENABLE_MPS_FALLBACK=1 enabled, it kept crashing for the longer strings.1KPipeline benchmark for voice af_heart (warm-up took 0.568s) using pip package
2Test Chars Output (s) Inf(s) RTFx Peak GB
31 42 2.750 0.414 6.649x 1.41
42 129 8.625 0.729 11.839x 1.54
5Total - 11.375 1.142 9.960x 1.54 1TTS benchmark for voice af_heart (warm-up took an extra 2.155s) using model prince-canuma/Kokoro-82M
2Test Chars Output (s) Inf(s) RTFx Peak GB
31 42 2.750 0.347 7.932x 1.12
42 129 8.650 0.597 14.497x 2.47
53 254 15.525 0.825 18.829x 2.65
64 93 6.125 0.306 20.039x 2.65
75 104 7.200 0.343 21.001x 2.65
86 130 9.300 0.560 16.611x 2.65
97 197 12.850 0.596 21.573x 2.65
108 6 1.350 0.364 3.706x 2.65
119 1228 76.200 2.979 25.583x 3.29
1210 567 35.200 1.374 25.615x 3.37
1311 4615 286.500 11.112 25.783x 3.37
14Total - 461.650 19.401 23.796x 3.37~15s to compile the model on the first run, subsequent runs are shorter, we expect ~2s to load.1> swift run fluidaudio tts --benchmark
2...
3FluidAudio TTS benchmark for voice af_heart (warm-up took an extra 2.348s)
4Test Chars Ouput (s) Inf(s) RTFx
51 42 2.825 0.440 6.424x
62 129 7.725 0.594 13.014x
73 254 13.400 0.776 17.278x
84 93 5.875 0.587 10.005x
95 104 6.675 0.613 10.889x
106 130 8.075 0.621 13.008x
117 197 10.650 0.627 16.983x
128 6 0.825 0.360 2.290x
139 1228 67.625 2.362 28.625x
1410 567 33.025 1.341 24.619x
1511 4269 247.600 9.087 27.248x
16Total - 404.300 17.408 23.225
17
18Peak memory usage (process-wide): 1.503 GB