This dataset was automatically built using a custom pipeline to format data for Sentence Transformers training.
Original Dataset: sentence-transformers/natural-questions
Split processed: train
Anchor Column: query
Positive Column: answer
Negative Column: negative
Hard negatives were aggressively mined from the dataset to improve model training robustness.