ODA-Mixture-100k is a compact general-purpose post-training dataset curated from top-performing open corpora (selected via the OpenDataArena leaderboard) and refined through deduplication, benchmark decontamination.
Domain: General-purpose(e.g., Math, Code, Reasoning, General).
Format: Problem → Solution (reasoning trace) → Final answer.
Scale (selected training set): ~100K samples.
Goal: Achieve significant general-purpose… See the full description on the dataset page:
https://huggingface.co/datasets/kunato/ODA-Mixture-100k.