A small, curated benchmark for clustering documents by their content topic
(not by their visual form/layout). Each item is a single document page provided
as an image plus two text views (a VLM description and OCR markdown), with a
ground-truth content class.
The set is intentionally "tangle-stripped": 27 borderline items whose content
sits ambiguously between two classes were removed from a larger 907-item pool to… See the full description on the dataset page:
https://huggingface.co/datasets/langminer/doc-content-clustering-740.