This model is a fine-tuned version of Llama-3.1-8B-Instruct on the
sciworld-metaplan-preference-pairs dataset.
It achieves the following results on the evaluation set:
This model uses Meta Plan Optimization (MPO) to improve the planning capabilities of LLM agents. It leverages high-level general guidance through meta plans and enables continuous optimization based on feedback from the agent's task execution. It achieves state-of-the-art performance on ALFWorld and SciWorld, with an average accuracy of 83.1.
The model was trained on the
sciworld-metaplan-preference-pairs dataset, part of the
Meta_Plan_Optimization dataset.