Text embeddings for ~200,000 Wikidata entities, generated from each entity's
English label with Google's open google/embeddinggemma-300m model. Vectors
are 768-dimensional and unit length (the model normalizes its output), so
cosine similarity equals a plain dot product.
entity_id
string
Wikidata Q-identifier, e.g. Q42
label
string
English label, e.g. Douglas Adams… See the full description on the dataset page:
https://huggingface.co/datasets/dejanseo/200k-wiki-data-embeddings-gemma.