A curated dataset of 100 high-quality prompts designed for benchmarking Large Language Model (LLM) performance across various metrics including latency, throughput, and response quality. This dataset features very long, complex prompts ideal for testing models' capabilities with extended context, creative writing, and detailed narrative generation.
Size: 100 prompts
Format: JSONL (JSON Lines)
Average Token Length: Variable (very… See the full description on the dataset page:
https://huggingface.co/datasets/raffel36/benchmark_16k.