Companion training dataset for the KnowRL projectPaper: KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance (2026)
KnowRL-Train-Data is the training corpus used for KnowRL-style RLVR optimization with minimal-sufficient knowledge-point guidance.
This dataset includes:
Original problem source metadata (data_source)
Rule-based reward supervision (reward_model)
Final training prompt in chat… See the full description on the dataset page:
https://huggingface.co/datasets/HasuerYu/KnowRL-Train-Data.