Recent advances in agent frameworks have pushed large language models beyond conversational assistance toward goal-driven task execution. Among these frameworks, OpenClaw has emerged as a representative setting for evaluating whether models can interact with tools, follow multi-step instructions, and complete practical tasks in realistic environments. Unlike traditional chatbot benchmarks, OpenClaw-style scenarios require models not only to produce… See the full description on the dataset page:
https://huggingface.co/datasets/whiskey1983/ZClawBench.