A 2.5 GB CSV corpus teaching LLMs to minimize token usage in their outputs.
Progresses from basic filler removal to expert-level nested reasoning compression.
verbose_output - The padded, wasteful version of the text
efficient_output - The compressed, token-efficient equivalent
technique - Compression strategy used
subcategory - Specific variant of the technique
difficulty - Tier 1 (easiest)… See the full description on the dataset page:
https://huggingface.co/datasets/Gugu8/Token-Efficiency.