This is a sampled big_patent dataset - sampled down for shorter fine-tunings.
The data is sampled with the aim of providing an even distribution across data lengths. The distribution is quite flat up until 1 million characters in length, making the dataset good for training on lengths up to 250,000 tokens.