STAR-1 is a high-quality safety dataset designed to enhance safety alignment in large reasoning models (LRMs) like DeepSeek-R1.
Built on the principles of diversity, deliberative reasoning, and rigorous filtering, STAR-1 integrates and refines data from multiple sources to provide policy-grounded reasoning samples.
The dataset contains 1⦠See the full description on the dataset page:
https://huggingface.co/datasets/UCSC-VLAA/STAR-41K.