T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
✨ Introduction
This is an evaluation harness for the benchmark described in T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step.
[Paper]
[Project Page]
[LeaderBoard]
[HuggingFace]
Large language models (LLM) have achieved remarkable performance on various NLP tasks and are augmented by tools for broader applications. Yet, how to evaluate and… See the full description on the dataset page: https://huggingface.co/datasets/lovesnowbest/T-Eval.