This dataset accompanies the paper A Framework for Evaluating Agentic Skills
at Scale and supports rigorous, reusable evaluation of agent skills —
structured, reusable knowledge artifacts that augment LLM agent capabilities.
Each task is a realistic, end-to-end challenge derived from a real-world skill
(or a group of related skills) sourced from open-source GitHub repositories of
trusted, well-known organizations. A… See the full description on the dataset page: https://huggingface.co/datasets/tesslio/task-evals-for-skills.