This dataset is a derivative work of MIRACL (Zhang et al., 2023)
restricted to the Indonesian (id) subset, preprocessed five different ways
to study how Indonesian -nya clitic handling affects retrieval quality.
Licensed under Apache-2.0, matching MIRACL.
Preprocessing strategies
keep
Baseline pass-through. Text is preserved exactly as MIRACL ships it.