Multi-model embedding dataset built on HuggingFace FineWiki, designed for approximate nearest neighbor (ANN) search benchmarking with USearch and other vector search engines.
The same Wikipedia corpus — chunked, cleaned, and enriched with graph metadata — is embedded by multiple models spanning dense BERT-like encoders, GPT-style decoder-based LLMs, and late-interaction ColBERT-style architectures.
Each model's embeddings ship with precomputed ground-truth k-nearest… See the full description on the dataset page:
https://huggingface.co/datasets/unum-cloud/USearchWiki.