This repository contains unified datasets used by the LLM-planning framework.
The files here are organized to match the current experiment entrypoint in scripts/exp.sh and the multi-stage planning pipeline used in the repo.
Augmented GAIA
4 category folders + DAG/reference folders
165 main eval samples
Multimodal answer-based benchmark with attachments, GPT-4o dependency… See the full description on the dataset page:
https://huggingface.co/datasets/Alfiechuang/llm.planning.