π arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
π€ Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning.
SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline.
However, it had two significant limitations:
(1) it was distributed under the Llama 3.3 Community License, and
(2) its size was limited to⦠See the full description on the dataset page:
https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.