description: |
This dataset contains scraped data from IslamQA.
It includes the original questions and embeddings of the answers.
Various metadata scraped from the website, such as the category of the question, are also included.
The following two models were used to generate embeddings:
We provide embeddings of the truncated answer (to fit model limits) and the embeddings of every sentence. Sentences are tokenized using spaCy.… See the full description on the dataset page:
https://huggingface.co/datasets/IslamQA/islamqa.