Unlocking Data Value in Finance: A Study on Distillation
and Difficulty-Aware Training
ODA-Fin-SFT-318K is a meticulously curated financial reasoning dataset comprising 318,599 samples with high-quality Chain-of-Thought (CoT) annotations. Constructed via multi-stage distillation from Qwen3-235B-A22B-Thinking and rigorous verification, this dataset establishes a robust foundation for training financial language models with strong reasoning capabilities.⦠See the full description on the dataset page:
https://huggingface.co/datasets/OpenDataArena/ODA-Fin-SFT-318k.