The central artifact is a
seven-source hard-negative mixture mined with a multi-vector late-interaction model, rather
than a dense bi-encoder or BM25.
Candidate documents are retrieved under MaxSim: each query token is matched to its best
document token, and those best similarities are summed. This produces negatives that are hard
under the same token-aware scoring geometry GLINT uses at retrieval time.
Multi-vector mining model:… See the full description on the dataset page:
https://huggingface.co/datasets/chungimungi/GLINT-data.