MiniLM-L12 cross-encoder reranker trained under the configuration stated in §4 of
Askari et al., "Injecting the BM25 Score as Text Improves BERT-Based Re-rankers", ECIR 2023:
cross-entropy loss with
early stopping on validation nDCG@10.
Paired with
Amdestya/ce-bm25cat-distilbert — identical data,
schedule and seed; the arms differ only in the input construction.
1text_a = f"{query}"
2text_b = passage
3# tokenises to: [CLS] query [SEP] passage [SEP]
Matched no-injection baseline. Do NOT inject anything into text_a -- this model has never seen a numeric token there.
1from sentence_transformers import CrossEncoder
2model = CrossEncoder("Amdestya/ce-cat-distilbert", max_length=256)
Outputs are raw logits (identity activation), unbounded, higher = more relevant.
Built as a reference artifact for a dissertation study on reproducing IR papers with LLMs.