One format. Clear provenance. Parser-verified. The current snapshot has 15 evaluation subsets, ready to eval.
Most math datasets on HF come in different shapes — Parquet, JSON, tar.gz, custom splits, inconsistent answer formats. math-vault normalizes them all to a single canonical schema (id / problem / answer / dataset) and verifies every answer through a unified parser.
This repo is the consumption layer. For full source history, per-dataset provenance READMEs… See the full description on the dataset page:
https://huggingface.co/datasets/Geraldxm/math-vault.