12 public example tasks from ASCIITermDraw-Bench, a benchmark for
evaluating whether language models can generate and edit structured ASCII
diagrams.
The full benchmark has 80 private, held-out tasks used for actual scoring —
those are not distributed here. This dataset is a separate, hand-authored
set of 12 tasks (one easy, one medium, one hard per category) in the
exact same format, so anyone can see what a task looks like and run the… See the full description on the dataset page:
https://huggingface.co/datasets/YuvrajSingh9886/asciitermdraw-bench-public.