Supervised finetuning data for RecaLLM. Contains 1,795 reasoning traces annotated with verbatim recall spans, used as a cold start before GRPO reinforcement learning.
Six teacher models (Llama-70B-R1, Llama-8B-R1, Qwen3-32B, Qwen3-8B, Qwen3-30B-A3B, Qwen3-8B-R1) generated reasoning traces across four retrieval and reasoning tasks.
Only traces that produced the correct final answer were retained.
GPT-5.2 rewrote… See the full description on the dataset page:
https://huggingface.co/datasets/kswhitecross/RecaLLM-sft.