Effective evaluation of multi-hop tool use is critical for analyzing the understanding, reasoning, and function-calling capabilities of large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/bytedance-research/ToolHop.