A lightweight dataset of (anchor, positive, negative) triplets for training Vietnamese medical text embeddings.
Anchors are short section headers, positives are answer snippets from the same article, and negatives are semantically related snippets from other articles in the same category (semi-hard negatives).
Language: Vietnamese
Domain: Healthcare / Patient education
Format: JSON
Use cases: Contrastive learning (TripletLoss /… See the full description on the dataset page:
https://huggingface.co/datasets/mtue29/vietnamese-medical-dataset.