Real Bangla and English government forms assembled into document packets for the
packet-splitting task: given a packet of concatenated form pages, recover which pages
belong to each source document and restore each document's original page order.
This is the dataset and its schema. The full benchmark pipeline (inference,
evaluation, and analysis) lives in the code repository:… See the full description on the dataset page:
https://huggingface.co/datasets/Mausul/khondo.