Agent traces from PostTrainBench (GitHub), a benchmark that measures CLI agents' ability to post-train base LLMs.
A pre-trained base LLM to fine-tune
An evaluation script for a specific benchmark
10 hours on an NVIDIA H100 80GB GPU
The agent must autonomously improve the model's performance on the target benchmark using any post-training strategy it chooses (SFT, LoRA, RLHF, prompt engineering for data… See the full description on the dataset page:
https://huggingface.co/datasets/aisa-group/PostTrainBench-Trajectories.