Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page:
https://huggingface.co/datasets/mannycooper/document-review-data.