Synthetic training and evaluation data for fine-tuning embedding and reranking models on the Snap / Ubuntu Core / SnapD domain. Generated from Canonical documentation using Gemini 2.5 Flash via Vertex AI.
Anchor/positive pairs for embedding model training (MultipleNegativesRankingLoss).Columns: anchor, positive
Preference triplets for DPO fine-tuning of causal reranking… See the full description on the dataset page:
https://huggingface.co/datasets/Rnfudge/snapd-training-data.