This dataset is a collection of title-body pairs collected from NPR using Pushshift. See embedding-training-data for additional information.
This dataset can be used directly with Sentence Transformers to train embedding models.
Dataset Subsets
pair subset
Columns: "title", "body"
Column types: str, str
Examples:{
'title': "Al-Awlaki's Death Raises Questions About U.S. Tactics",
'body': "A joint CIA and U.S. military operation targeted… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/npr.