실제 embedding 학습 queue가 소비하는 provenance-preserving 파생 JSONL이다.
source-native query/positive 관계와 deterministic bootstrap negative로 컴파일했다. 최종 학습 전 current-student hard-negative mining을 수행해야 한다.
rows: 250,000
release eligible: true
visibility: public
use: public redistribution and model training
exact benchmark query/evaluation-text matches: 0
exact retrieval-corpus matches: 0 unique hashes
Sionic 9 및 MTEB task-family train source가 포함될 수… See the full description on the dataset page:
https://huggingface.co/datasets/LLM-OS-Models2/ko-legal-embedding-training-v1.