This dataset contains the Stage 2 distillation data used to train
ZipRerank, a framework for highly efficient
list-wise multimodal rerankers for long documents.
The relevance rankings were generated by prompting GPT-5-mini to rerank
first-stage candidate pages from the
MMDocIR training set. These GPT-generated
rankings act as soft targets for knowledge distillation during ZipRerank's Stage 2
fine-tuning (Soft Ranking… See the full description on the dataset page:
https://huggingface.co/datasets/dukesunmtri/ZipRerank_GPT-5-mini_MMDocIR_Train.