The PIPKIN 174K Creative dataset is the largest dataset in the series of evolving, in-the-wild creative datasets.
This dataset is inspired by Pygmalion's PIPPA dataset from 2023.
Data is collected by exchanging anonymous data for OSS model usage (GLM-5, GLM 4.7, DeepSeek V3.1, Qwen 3.5 397B, Kimi K2.5, etc).
This data shows 1.6 billion tokens of chat data.
The average amount of input tokens is ~1.2k.
The average amount of completion tokens is ~1235.
There are… See the full description on the dataset page:
https://huggingface.co/datasets/LucidityAI/PIPKIN-Creative-174k.