Synthetic parallel-language data generated from paired synchronous
context-free grammars (SCFGs). Each experiment config holds three splits:
grammars — one row per grammar, with the SCFG source, lexicons, and
typological knobs for both language a and language b.
samples — generated (left, right) sentence pairs joined to a grammar by
grammar_name.
shots — held-out example pools for few-shot prompting. This split is empty
for experiments that do not define shots_*.jsonl… See the full description on the dataset page:
https://huggingface.co/datasets/jowenpetty/scfg.