PyBench: Evaluate LLM Agent on Real World Tasks
📃 Paper
•
🤗 Data (PyInstruct)
•
🤗 Model (PyLlama3)
•
Code
•
PyBench is a comprehensive benchmark evaluating LLM on real-world coding tasks including chart analysis, text analysis, image/ audio editing, complex math and software/website development. We collect files from Kaggle, arXiv, and other sources and automatically generate queries according to the type and content of each file.
The LLM Agent, equipped… See the full description on the dataset page:
https://huggingface.co/datasets/Mercury7353/PyInstruct.