๐ arXiv: Read our paper for detailed methodology at arXiv:2505.02881.
๐ค Sister Dataset: Discover SwallowCode2, our companion dataset for code generation.
SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1.
Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open andโฆ See the full description on the dataset page:
https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.