The PIPKIN 30K Creative dataset is the first dataset in a series of evolving, in-the-wild creative datasets.
This dataset is inspired by Pygmalion's PIPPA dataset from 2023.
Data is collected by exchanging anonymous data for OSS model usage (GLM-5, GLM 4.7, DeepSeek V3.1, Qwen 3.5 397B, Kimi K2.5, etc).
This data shows 200+ million tokens of chat data.
The average amount of input tokens is ~10k.
The average amount of completion tokens is ~0.8k-1k.
There are 28k… See the full description on the dataset page:
https://huggingface.co/datasets/LucidityAI/PIPKIN-30K-Creative.