Synthetic English pretraining corpora generated from a Penn Treebank-derived constituency PCFG.
This release contains three corpus variants with different subject-verb agreement cue reliability settings:
The corpora use opaque token IDs rather than English word forms and were created for experiments on structural transfer from… See the full description on the dataset page:
https://huggingface.co/datasets/lluten/ptb_pcfg.