An OCR-degraded version of the MIRACL multilingual retrieval benchmark (miracl/miracl), designed to evaluate embedding models on noisy, OCR-like text.
A 2,000-document subsample per language was drawn from miracl/miracl (dev split).
Each passage and query was rendered as a PDF at a specific DPI / font-size setting and
re-extracted via OCR using the
ocr-robust-multilingual-embeddings
OCR simulator to introduce realistic character-level noise.… See the full description on the dataset page:
https://huggingface.co/datasets/Psychias/ocr-miracl.