📚 Our Paper (EMNLP 24 Resource Award)
A User-Centric Multi-Intent Benchmark for Evaluating Large Language Models.
User Reported Scenario (URS) Dataset
Dataset Features
Real-world usage scenarios of LLMs
The dataset is collected through a User Survey with 712 participants from 23 countries in 6 continents.
System abilities and performances in different scenarios might be different
Users’… See the full description on the dataset page:
https://huggingface.co/datasets/JiayinWang/URS.