Curated mirror of openbmb/Ultra-FineWeb maintained by the majentik MLX-quantization project.
Bilingual pretraining-corpus calibration subset for quantizing the MiniCPM5 family and other OpenBMB-lineage models (the corpus these models were trained on). Complements majentik/c4-calib with in-distribution and Chinese text.
4096 en + 1024 zh docs, streamed… See the full description on the dataset page:
https://huggingface.co/datasets/majentik/ultra-fineweb-calib.