The dataset used in Black-Box On-Policy Distillation of Large Language Models paper. Homepage at here.
This dataset is an extension of the LMSYS-Chat-1M-Clean corpus, specifically curated by collecting high-quality, non-refusal responses from the GPT-5-Chat API.
The LMSYS-Chat-1M dataset collects real-world user queries from the Chatbot Arena.
There is no tool calls or reasoning in the GPT-5-Chat response.
The… See the full description on the dataset page:
https://huggingface.co/datasets/Rafi757/Frontier.