Collection of conversations evaluated using Qwen 3 series.
Prompt template:
You are an AI evaluator tasked with rating the overall difficulty of a complete human–AI conversation (all user messages and AI responses) on a 1–10 scale based on how challenging it would be for an AI to handle effectively.
[CONVERSATION]
Evaluate the conversation as a whole, considering:
- Clarity of user intent
- Required context and reliance on prior turns… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/chat-difficulty.