This dataset contains OCR results from images in stephenmcconnachie/0004-pdf-pages-test using Falcon OCR, a 0.3B early-fusion vision-language model.
Source Dataset: stephenmcconnachie/0004-pdf-pages-test
Model: tiiuae/Falcon-OCR
Task Mode: plain - Full-page text extraction
Number of Samples: 50
Processing Time: 1.8 min
Processing Date: 2026-06-20 09:46 UTC
Backend: falcon-perception… See the full description on the dataset page:
https://huggingface.co/datasets/stephenmcconnachie/0004-pdf-falcon-ocr.