This repository contains the accepted EDGAR-OCR benchmark samples used to evaluate OCR-style transcription of complex SEC filing tables.
samples/<sample_id>/screenshot.png: rendered table image used as model input.
samples/<sample_id>/synthetic_table.html: synthetic filing-style HTML table used to render the input.
samples/<sample_id>/ground_truth_table.md: target table transcription.
samples/<sample_id>/ground_truth_grid.json: parsed target… See the full description on the dataset page:
https://huggingface.co/datasets/sfd-anonymous/edgar-ocr-benchmark.