A controlled ambiguity benchmark for exact CFG/PCFG parsing.
The dataset is built for the paper
Symplectic Inside--Outside Atlas for Ambiguous Grammars
and is intended for chart-level diagnostics, ambiguity stress tests,
weighted deduction experiments, and parser-evaluation studies.
The main point is simple: parse count alone does not determine
ambiguity geometry. Two examples can have the same number of parses
while concentrating uncertainty in… See the full description on the dataset page:
https://huggingface.co/datasets/Lightcap/symplectic-inside-outside-atlas.