This dataset provides long prompts intended to use for testing language models with long inputs with different sizes.
id: numerical id
source: the source of the prompt text
length: indication of length of the prompt
text: the text of the prompt
Exact prompt length depend on the tokenizer for the model, and can differ quite a bit with different tokenizers.
The lengths included in the dataset should be seen as… See the full description on the dataset page:
https://huggingface.co/datasets/helenai/summarize-long-texts.