Korean financial-domain (anchor, positive) pair dataset for sentence-embedding model training and as a retrieval corpus for evaluation.
This dataset acts in two roles:
Training input - 45,589 (anchor, positive) pairs for in-batch / hard-negative contrastive learning.
Retrieval corpus - source_chunk_id (≈36k unique chunks after q-variant dedupe) is the candidate pool for IR evaluation in the companion BCCard/BCAI-Finance-Kor-Embedding-Triplet dataset.… See the full description on the dataset page:
https://huggingface.co/datasets/BCCard/BCAI-Finance-Kor-Embedding-Pair.