The dataset is a random sample of Historical English Books title-abstract pairs and query-sentences pairs. The data is primarily from the Victorian Era up until 1899.
The title and abstracts are used as is from the upstream datasource. Each book is split by sentence and a random sample of sentences are selected per split. Logic was used to validate there is no overlap between the splits. Queries are LLM-generated using the… See the full description on the dataset page:
https://huggingface.co/datasets/NeuML/historical-english-books-similarity.