A repository-level code retrieval dataset built from what real developers actually do together. Positives are labeled by co-change — files that engineers repeatedly modified in the same commits — not by imports, folder co-location, or any other structural proxy. Every training row ships with 128 hard + 128 easy negatives, where hard negatives are mined in the embedding space of a strong open model (Qwen3-Embedding-4B) and filtered through an eight-signal… See the full description on the dataset page:
https://huggingface.co/datasets/aysinghal/cochange-code-retrieval.