The full judge-alignment dataset from CuBE (Culturally-situated Behavioral Evaluations): culturally-localized vignettes paired with both crowdsourced human labels and parallel LLM-judge labels, used to measure how well off-the-shelf LLM judges agree with culturally-situated human annotators on subjective behavioral interpretation.
For each of 11 alignment questions spanning 6 risky behaviors, a culturally-localized vignette… See the full description on the dataset page:
https://huggingface.co/datasets/token-opt-org/cube-judge-alignment.