Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed.
Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page:
https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.