FlexBench is a benchmark for evaluating strictness-adaptive content moderation under policy shifts. Each sample is annotated with a 5-tier risk severity label (BENIGN / LOW / MODERATE / HIGH / EXTREME). Following the accompanying paper, we derive three deployment-oriented binary classification tasks—strict, moderate, and loose—by… See the full description on the dataset page: https://huggingface.co/datasets/Tommy-DING/FlexBench.