Synthetic bracket-tagged coreference data for training multilingual
coref pointer-network models. Generated with Claude (Haiku 4.5 on
Bedrock) via the templated prompts in
infon.cassette.synthgen_coref.
Each row:
{
"doc_id": "coref_synth_
",
"lang": "en|ja|zh|ko|th",
"domain": "automotive|defence|semiconductors|energy",
"text": "...",
"mentions": [
{"start": 0, "end": 6, "text": "Toyota", "cluster":… See the full description on the dataset page: https://huggingface.co/datasets/cp500/infon-coref-multilingual.