This dataset is a collection of question-answer pairs, collected from Google. See GooAQ for additional information.
This dataset can be used directly with Sentence Transformers to train embedding models. This dataset is equivalent to sentence-transformers/gooaq, but with 2500 evaluation samples and 2500 test samples randomly extracted from the training set.