This dataset card provides information for the dataset attached to the AI Fiction in the Wild(chat) paper. The dataset presentes over 500K English Wildchat conversations that have been labelled by an LLM across three axis, fictional, fanfiction, and sexually explicit. The goal of this project is to understand the types of fiction being generated by real LLM users.
We use the original Wildchat dataset, before filtering down to… See the full description on the dataset page:
https://huggingface.co/datasets/neelgupta2112/Wildchat-1M-English-Fiction-Labels.