This dataset is a preprocessed version of Eureka-Lab/PHYBench, containing only the samples that have a complete solution and answer.
The data has been formatted into a question and answer structure suitable for training instruction-following language models.
question: The original physics problem statement (from the content column).
answer: A string containing the thinking process and the final answer, formatted as… See the full description on the dataset page:
https://huggingface.co/datasets/LLMcompe-Team-Watanabe/PHYBench_preprocess.