In this dataset, ~32k human responses collected in less than 1h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
The annotators were asked Which voice is more friendly? and Which voice sounds more natural? respectively.