GAND (Gender-Ambiguous Natural Data) is a benchmarking resource for evaluating gender (bias) in machine translation or downstream NLP tasks.
The data stems purely from natural data resources (OpenSubtitles from the OPUS project and C4).
The data has been meticulously (automatically + manually) filtered to ensure complete gender ambiguity with respect to a specific referent.
More information on the compilation of GAND can be found on GitHub.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/jhacken/GAND.