Measuring whether agents can improve reusable skills, and whether those
improvements transfer across roles, tasks, and execution contexts.
📄 Abstract
AFTER is a benchmark for studying skill evolution: the ability of an
agentic framework to revise, specialize, and reuse skill instructions after
observing task experience. Unlike task-only evaluation, AFTER separates the
problem into… See the full description on the dataset page: https://huggingface.co/datasets/DavydenkoGr/AFTER.