A diverse collection of synthetic enterprise documents for benchmarking context extraction and RAG systems.
This dataset contains 16 representative enterprise documents spanning multiple formats and domains, designed to evaluate:
Structure-aware indexing - Can the system identify high-value vs. low-value content?
Time decay relevance - Does the system properly weight recent vs. old information?
Pragmatic truth detection - Can… See the full description on the dataset page:
https://huggingface.co/datasets/imran-siddique/context-as-a-service.