The SemBenchmarkClassificationSorted dataset is a streamlined and sorted variant of SemBenchmarkClassification, designed for sequential processing and efficient evaluation of semantic caching systems in non-i.i.d. tasks.
This dataset is derived from the original SemBenchmarkClassification benchmark with the following modifications:
Response-Based Sorting: All 45,000 examples are sorted by their… See the full description on the dataset page:
https://huggingface.co/datasets/vCache/SemBenchmarkClassificationSorted.