Text quality prediction (BLEU) for various document parsers given the first page's PyMuPDF-extracted text.
This model relates to the second version of AdaParse ("AdaParse v2").
Model Details
Allen AI's Specter fine-tuned for document quality prediction.
Predict quality of parser output given the extracted text/
Direct Use
Document quality prediction for resource-optimal delegation within AdaParse (version 2 for this particular instance).
Downstream Use [optional]
[More Information Needed]
Out-of-Scope Use
Quality prediction for documents that are (a.) out-of-distribution (e.g., non-scientific) or (b.) for parsers that were not part of the fine-tunign set.
Bias, Risks, and Limitations
Bias: Model was trained on tens of thousands of scientific documents from several journals across eight scientific disciplines (mathematics, engineering, biology, physics, etc.). Naturally, biased towards STEM documents.
Limitations: Quality prediction based on a single page's text of one particular extraction tool (PyMuPDF) is challenging.
High-performance compute (Aurora, Polaris, Sophia, Lambda) at Argonne National Laboratory (ANL)/Argonne Leadership Computing Facility (ALCF).
Citation
BibTeX:
@article{siebenschuh2025adaparse,
title={AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine},
author={Siebenschuh, Carlo and Hippe, Kyle and Gokdemir, Ozan and Brace, Alexander and Khan, Arham and Hossain, Khalid and Babuji, Yadu and Chia, Nicholas and Vishwanath, Venkatram and Stevens, Rick and others},
journal={arXiv preprint arXiv:2505.01435},
year={2025}
}
APA:
Siebenschuh, C., Hippe, K., Gokdemir, O., Brace, A., Khan, A., Hossain, K., ... & Underwood, R. (2025). AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine. arXiv preprint arXiv:2505.01435.