This dataset has been created with distilabel.
This dataset obtains embeddings for the dataset argilla-warehouse/personahub-fineweb-edu-4-dedup,
using the Alibaba-NLP/gte-large-en-v1.5 model from sentence transformers.
The pipeline can be seen at: pipe_personahub_embeddings.py.
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline… See the full description on the dataset page:
https://huggingface.co/datasets/argilla-warehouse/personahub-fineweb-edu-4-embeddings.