This repository contains MiniLLM_MultiLingual_V2 embeddings for the CrediBench WebContent (December 2024) dataset.
Source dataset:
https://huggingface.co/datasets/Hussein-Abdallah/CrediBench-WebContent-Dec2024_CochranSampled
The source dataset is a Cochran-sampled subset of the December 2024 CrediBench WebContent corpus. Documents are sampled independently for each web domain using Cochran's sampling… See the full description on the dataset page:
https://huggingface.co/datasets/credi-net/CDB_DEC2024-CochranSampled_MiniLLMV2_Emb.