ODA-Mixture-500k is a large-scale general-purpose post-training dataset curated from top-performing open corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination.
Domain: General-purpose(e.g., Math, Code, Reasoning, General).
Format: Problem → Solution (reasoning trace) → Final answer.
Scale (selected training set): ~500K samples.
Goal: Achieve maximum general-purpose… See the full description on the dataset page:
https://huggingface.co/datasets/szhyxt/turman.