A benchmark for measuring how well ML engineering agents handle ambiguous evaluation metrics in Kaggle-style competitions.
Each task is a Kaggle competition from MLE-bench (OpenAI, 2024). For every task we provide two prompt variants — one in which the true evaluation metric is named, and one in which it is redacted. The agent must produce a submission CSV that is graded against the true metric using MLE-bench's grading infrastructure.
The… See the full description on the dataset page:
https://huggingface.co/datasets/anonymous222bit/Ambig-DS-M.