O1aw-Dataset is a comprehensive legal question-thought-answer dataset, designed to evaluate and enhance legal reasoning capabilities in language models. The dataset follows the O1-style format, featuring complex legal scenarios that require multi-step reasoning.
First, we crawl and clean raw legal materials from the internet, including Hong Kong e-Legislation. Then, we use GPT-4o to generate corresponding questions and… See the full description on the dataset page:
https://huggingface.co/datasets/HKAIR-Lab/HK-O1aw-SFT-16K.