This dataset extends nlpai-lab/ko-triplet-v1.0 with freshly mined hard negatives for each (anchor, positive) pair, intended for fine-tuning Korean text embedding models.
The hard negatives were identified using Qwen/Qwen3-Embedding-8B, retrieving candidates from a corpus built from the deduplicated union of the original dataset's document and hard_negative columns.
positive (string): Document relevant to the… See the full description on the dataset page:
https://huggingface.co/datasets/whybe-choi/ko-triplet-hn.