PARROT is a large-scale benchmark designed to evaluate the ability of models to generate executable data preparation (DP) pipelines from natural language instructions. It introduces a new task that aims to lower the technical barrier of data preparation by translating human-written instructions into code. To reflect real-world usage, the benchmark includes ~18,000 pipelines spanning 16 core transformation operations, built from 23,009 tables across six public datasets.
This benchmark is… See the full description on the dataset page:
https://huggingface.co/datasets/momo006/PARROT.