Phinc is a parallel corpus for machine translation pairing code-mixed Hinglish (a fusion of Hindi and English commonly used in modern India) with human-generated English translations.
You can evaluate an embedding model on this dataset using the following code:… See the full description on the dataset page:
https://huggingface.co/datasets/mteb/PhincBitextMining.