DocBlocks is a high-quality, multilingual document-level machine translation (MT) dataset designed to fine-tune large language models (LLMs) on long-context translation tasks. Unlike traditional sentence-level datasets, it contains full documents with natural discourse structures and contextual alignment, helping models maintain coherence, consistency, and high translation quality across longer texts.
Curated by: Instituto Superior Técnico, Instituto de… See the full description on the dataset page:
https://huggingface.co/datasets/sardinelab/DocBlocks.