FineWeb2-embeddings is an extension of the FineWeb2 dataset, annotated with document-level Snowflake's Arctic-embed-m-v2.0 embeddings for 36 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research.
Snowflake-arctic-embed-m-v2.0 has a sequence length limit of 8192 tokens, each document's embeddings are obtained by using the CLS token to embed each document.… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_embeddings.