Independent AI safety research lab specializing in cognitive fit, alignment, and human-AI collaboration
A curated dataset of 1,400 conversational examples demonstrating how to decline unhelpful, misguided, or counterproductive requests while explaining the reasoning and offering constructive alternatives. Designed for fine-tuning language models to be genuinely helpful by knowing when and how to say no.… See the full description on the dataset page:
https://huggingface.co/datasets/vanta-research/reasoned-refusal.