Pre-tokenized screenplay corpus used to train and evaluate the models in the GPT-2 Screenplay Fine-Tuning Study. Derived from Movie-Script-Database by Aveek Saha. Provided as tokenized JSON splits ready for direct consumption by a GPT-2 Trainer pipeline — no preprocessing required.
This dataset contains approximately 94 million tokens of professionally formatted screenplay text, pre-tokenized using the… See the full description on the dataset page:
https://huggingface.co/datasets/kazkiryuu/Films.