This public release directly appends 10,284 accepted legacy title-generation
rows and 13,500 nine-language LLM-generated rows. The 23,784 examples are split
as train 20,584, validation 1,550, legacy test 200, legacy Vietnamese test 100,
and label-free synthetic holdout 1,350. Legacy rows contain only messages;
nine-language rows retain their richer IDs, language, coverage, cluster,
quality, and model-provenance fields. Train and validation… See the full description on the dataset page:
https://huggingface.co/datasets/ManhHoDinh/titlegen-conversations.