Cleaned, chunked, and phase-organized training corpus for a small bilingual LLM
(Arabic + English) focused on Egyptian historical figures.
pretrain
phase1_train
175,476
Phase 1
General language competence (full phase_1, AR+EN)
pretrain
phase1_eval
2,000
Phase 1
Held-out perplexity eval (leakage-safe by article)
pretrain
phase2_train
94,486
Phase 2
Egypt-domain focus (use… See the full description on the dataset page:
https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset.