A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents"
Paper:
https://arxiv.org/abs/2406.13923
This dataset contains 14M samples in PIN format, with around 18.79 TB storage.
🚀 News
[ 2025.09.04 ] !NEW! 🔥 We have completed the final version of the PIN-14M dataset and conducted some simple statistics on it.
[ 2024.12.12 ] !NEW! 🔥 We have updated the quality signals for all subsets, with the dataset now containing 7.33B tokens… See the full description on the dataset page:
https://huggingface.co/datasets/m-a-p/PIN-14M.