HPLT2-embeddings is an extension of the HPLT2 dataset, annotated with document-level Snowflake's Arctic-embed-m-v2.0 embeddings for 35 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research.
Snowflake-arctic-embed-m-v2.0 has a sequence length limit of 8192 tokens, each document's embeddings are obtained by using the CLS token to embed each document.
The… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_embeddings.