DOAB Open Access Books - Metadata Extraction Dataset
Dataset Description
This dataset contains 9,363 open access books with page images and rich bibliographic metadata extracted from MARC21 records, curated specifically for training and evaluating Vision Language Models (VLMs) on automatic metadata extraction from scholarly monographs.
The dataset is derived from the Penn State ScholarSphere DOAB collection (Directory of Open Access Books), focusing on books with Creative… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/doab-metadata-extraction.