Repository:
https://anonymous.4open.science/r/JuICE
HuggingFace: juice-cultural-eval/JuiCE
We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh), in both… See the full description on the dataset page:
https://huggingface.co/datasets/juice-cultural-eval/JuICE.