Views
No views yet
SFT.py recipe, same official evaluation harness. It is an
internal research checkpoint for a controlled post-training evaluation.task_type=reasoning, English (42086 examples)## Thinking\n\n<reasoning>\n\n## Final Answer\n\n<answer>## Thinking then answers under ## Final Answer.eval.py + scorer.py, strict prompt),
so scores are drop-in comparable to MedReason paper Table 4.1@misc{wu2025medreason, title={MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs},
2 author={Juncheng Wu and others}, year={2025}, eprint={2504.00993}, archivePrefix={arXiv}}