A contrastive learning dataset for training cross-domain research paper retrieval models.
The goal is to embed papers so that those solving the same abstract problem cluster together — regardless of field, domain vocabulary, or application area.
For example, a paper on long-range dependency modeling in EEG signals and one on attention mechanisms for long sequences in NLP should be near each other in embedding space, even though they share no surface… See the full description on the dataset page:
https://huggingface.co/datasets/Pravallika6/cross_domain_embeddings.