ETI Embedding Training Data — v4 (5-style, LLM-judged triplets)
Norwegian (nb) retrieval triplets built from the
NorskHelsenett/LOS_Document_classification_ETI
corpus of public-service / welfare / health documents.
Each row is a triplet (anchor, positive, negative) plus three metadata
columns (style, category, doc_url) you can use to filter, weight, or
build curriculum stages.
76,408 triplets · 2,530 source documents · 38,629 distinct anchors.