Views
No views yet
bigscience/bloom-1b7.
The LoRA adapter includes modules_to_save: ["classifier","score"], so the value head ships
with it.Meta-Okapi/<lang>_bloom1b7_judgerm_decay1e-6_lr5e-5_10ksteps models are full
fine-tunes on ~126,000 Okapi ranking pairs. This model is trained on a different, far
smaller corpus and should not be treated as a drop-in equivalent.| Okapi judges | this model | |
|---|---|---|
| training pairs | ~126,000 | 2,200 |
| steps x batch | 10,000 x 16 | 500 x 16 |
| epochs over data | ~1.3 | ~5.2 |
| adapter | full fine-tune | LoRA (r=16) |
nthakur/multilingual-ultrafeedback-binarized-dpo-v0.1), from a
strict id partition: the 2,200 pairs appear in no evaluation, adaptation or meta-training split
in any language.lr 1.67e-5 (constant), 500 steps, batch 16, weight_decay 1e-6, seq_length 1024, size_valid_set 0.30, eval split 300, early stopping patience 4 on eval_lossprompt + completion with truncation_side="right" and max_length=1024, matching
training. Left-truncation removes the prompt and collapses the model onto surface features.