An executable benchmark for end-to-end data-engineering agents in industrial environments.
DataClawEval measures an autonomous agent's ability to inspect data, implement and debug pipelines,
and materialize correct artifacts in realistic data-engineering workflows. It contains 100
production-grounded tasks across five execution engines: PySpark, MySQL, HiveSQL, PrestoSQL/Trino,
and FlinkSQL. Each task runs in an isolated Docker sandbox and is evaluated by a… See the full description on the dataset page:
https://huggingface.co/datasets/dicemy/DataClawEval.