Do LLMs adhere to the objective constraints of the situation's urgency, or do they default to learned preferences?
RoleConflictBench generates distinct expectations for two competing social roles held by the same person and synthesizes them into a first-person story depicting a role conflict. The benchmark evaluates how a model's decision changes depending on the situation's urgency, rather than defaulting to a fixed role preference.… See the full description on the dataset page:
https://huggingface.co/datasets/ddindidu/RoleConflictBench.