This dataset contains compressed versions of the Databricks Dolly-15k prompts. Each prompt was compressed using the gpt-5-nano model to minimize input tokens while preserving all constraints. You can explore the downstream model that relies on this data in the companion Space: Very Small Prompt Compression Demo.
Compression model: gpt-5-nano
Source dataset: databricks/databricks-dolly-15k
Rows: 15,000
Aggregate token savings: 289,540 → 215,219 tokens… See the full description on the dataset page:
https://huggingface.co/datasets/gravitee-io/dolly-15k-prompt-compression.