Views
No views yet
| Component | Specification |
|---|---|
| Architecture | Transformer decoder with RoPE, SwiGLU, RMSNorm, GQA |
| Layers | 24 |
| Hidden size | 896 |
| Attention heads | 14 (Q) / 2 (KV) |
| Original vocab | 151,665 |
| Pruned vocab | 30,556 |
| Expanded vocab | 30,557 (+1 special token </SEP>) |
| Embedding shape | (30,557, 896) |
| Tied embeddings | Yes (embed_tokens ↔ lm_head) |
</SEP> (ID: 30556) — Used as a document separator during packing.| Source | Type | Estimated Tokens |
|---|---|---|
| Shamela/Waqfeya | Classical Arabic books (tafsir, hadith, fiqh, aqidah, nahwu) | ~300 million tokens |
| Parquet shards | Pre-packed to 1,024 sequence length | 26,902 steps/epoch |
| Parameter | Value |
|---|---|
| Batch size per GPU | 16 |
| Gradient accumulation | 4 |
| Global batch (2× T4) | 128 |
| Learning rate (GaLore) | 1e-4 |
| Learning rate (embed/lm_head) | 5e-4 |
| LR scheduler | Cosine with warmup (200 steps) |
| Weight decay | 0.01 |
| GaLore rank | 128 |
| GaLore update gap | 400 |
| GaLore scale | 0.25 |
| Precision | FP16 (T4) |
| Gradient checkpointing | Enabled |
| Optimizer | GaLoreAdamW (3 parameter groups) |
torchrun --nproc_per_node=2| Metric | Base Model (Qwen2.5-0.5B) | This Model | Improvement |
|---|---|---|---|
| Avg Loss | 2.359 | 1.655 | ↓29.8% |
| Perplexity | 10.58 | 5.23 | ↓50.6% |
| Bits/Byte | 0.698 | 0.490 | ↓29.8% |
| Genre | Base PPL | Our PPL | Improvement |
|---|---|---|---|
| Quran | 3.36 | 1.69 | ↓49.6% |
| Hadith | 2.63 | 1.40 | ↓47.0% |
| Tafsir | 4.34 | 2.45 | ↓43.7% |
| Fiqh (Classical) | 5.72 | 2.41 | ↓57.8% |
| Prose Turath | 5.21 | 3.02 | ↓42.0% |
| Model | Distinct-1 | Distinct-2 | Distinct-3 |
|---|---|---|---|
| This model | 0.878 | 0.989 | 1.0 |
| Base | 0.880 | 0.957 | 1.0 |
Input: حَدَّثَنَا سُفْيَانُ عَنِ الزُّهْرِيِّ
Output: حَدَّثَنَا سُفْيَانُ عَنِ الزُّهْرِيِّ عَنْ عُبَيْدِ اللَّهِ بْنِ عَبْدِ اللَّهِ بْنِ عُتْبَةَ عَنْ أَبِيهِ قَالَ : كَانَ ر...
Input: وَالدَّلِيلُ عَلَى وُجُوبِ هَذِهِ الْمَسْأَلَةِ عِنْدَ الْفُقَهَاءِ
Output: وَالدَّلِيلُ عَلَى وُجُوبِ هَذِهِ الْمَسْأَلَةِ عِنْدَ الْفُقَهَاءِ أَنَّهَا إِذَا كَانَتْ مُحَرَّمَةٌ فَإِنَّهَا تَكُونُ مُحَرَّمَةٌ بِالْ...
Input: قَوْلُهُ تَعَالَى وَاصْبِرْ نَفْسَكَ مَعَ الَّذِينَ
Output: قَوْلُهُ تَعَالَى وَاصْبِرْ نَفْسَكَ مَعَ الَّذِينَ يَعْنِي بِأَرْحَامِهِمْ أَنفُسَهُمْ وَأَزْوَجَهُمْ وَذُرِّيَّتَهُمْ إِنَّ...
pip install transformers>=4.37.0 torch1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model_name = "Ik45/qwen2.5-0.5b-arabic-classic-pruned"
4
5model = AutoModelForCausalLM.from_pretrained(
6 model_name,
7 torch_dtype="auto",
8 device_map="auto",
9 trust_remote_code=True
10)
11tokenizer = AutoTokenizer.from_pretrained(
12 model_name,
13 trust_remote_code=True
14)
15
16# Generate
17prompt = "قَالَ الإمام مالك رحمه الله:"
18inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
19
20outputs = model.generate(
21 **inputs,
22 max_new_tokens=128,
23 do_sample=True,
24 temperature=0.7,
25 top_p=0.9
26)
27print(tokenizer.decode(outputs[0], skip_special_tokens=True))