A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents"
Paper:
https://arxiv.org/abs/2406.13923
This dataset contains around 200M samples in PIN format, with around 312 TB storage.
🚀 News
[ 2025.09.22 ] !NEW! 🔥 We have completed the final version of the PIN-200M dataset and conducted some simple statistics on it.
[ 2024.12.06 ] !NEW! 🔥 We have updated the quality signals, enabling a swift assessment of whether a sample meets… See the full description on the dataset page:
https://huggingface.co/datasets/m-a-p/PIN-200M.