Project Page | Paper
ReSA (Reasoned Safety Alignment) is an open-source synthetic safety-training dataset with 80K examples designed to enhance LLM robustness against jailbreak attacks through an "Answer-Then-Check" strategy. The dataset teaches models to first generate a summary of their intended answer, then critically evaluate its safety before providing a final response. This approach achieves superior safety performance while maintaining strong… See the full description on the dataset page:
https://huggingface.co/datasets/ByteDance-Seed/ReSA.