This dataset contains OCR results from images in markbaggett/amc-books using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance.
Source Dataset: markbaggett/amc-books
Model: zai-org/GLM-OCR
Task: text recognition
Number of Samples: 303
Processing Time: 34.2 min
Processing Date: 2026-06-23 12:40 UTC
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16… See the full description on the dataset page:
https://huggingface.co/datasets/markbaggett/amc-books-output.