710 Million (query, positive) pairs · 128 GB · 50+ Languages · 16 Sources
Pretraining data for Minnow-Em-v1 — part of KiteFishAI's Minnow family of sovereign small language models
This dataset is the stage-1 pretraining corpus used to train Minnow-Em-v1, KiteFishAI's multilingual embedding model and part of the Minnow family of sovereign small language models (SLMs). It aggregates 16 diverse open-source datasets into a… See the full description on the dataset page:
https://huggingface.co/datasets/KiteFishAI/kf-embed-pretrain-corpus-700M-raw.