Private dataset export for the canonical ds_eval_awareness_v2 coding benchmark from the non-verbal eval-awareness project.
samples.csv: 768 structured prompt rows
splits.csv: canonical split_eval_awareness_v2_all split file
summary.json: dataset summary and hashes
manifests/: provenance manifests used to build and govern the benchmark
256 BigCodeBench tasks
3 prompt conditions per task:… See the full description on the dataset page:
https://huggingface.co/datasets/Luxel/non-verbal-eval-awareness-benchmark-v2.