ALIGN is a benchmark for evaluating interaction-level alignment in generative AI systems through multi-turn interactions.
Traditional evaluation of AI systems focuses on output-level metrics (e.g., accuracy, correctness).
ALIGN shifts the focus toward interaction-level alignment, capturing how AI systems behave across multiple turns in real-world scenarios.
Scenario-based evaluation
Multi-turn… See the full description on the dataset page:
https://huggingface.co/datasets/Stellalala/ALIGN.