Views
No views yet
| Description | Value |
|---|---|
| Base Model | BSC-LT/MrBERT |
| Model Parameters | 308M |
| Tokenizer | SentencePiece (SPM) |
| Vocabulary Size | 256,000 |
| Hidden Layers | 22 |
| Hidden Size | 768 |
| Intermediate Size (FFN) | 1,152 |
| Attention Heads | 12 (head size: 64) |
| Attention Type | RoPE (Rotary Positional Embedding) |
| Sliding Window Size | 128 tokens |
| Global Attention Every N Layers | 3 |
| FFN Layer | Gated Linear Unit (GLU) |
| Normalization | Pre-norm LayerNorm (ε=1e-5) |
| Activation Function | GeLU |
| Precision | bfloat16 (AMP) |
| Context Length | 1,024 tokens |
| Weight Initialisation | Full Megatron |
| Hyperparameter | Value |
|---|---|
| Objective | Masked Language Modeling (MLM) |
| Train MLM Mask Probability | 30% |
| Eval MLM Mask Probability | 15% |
| Peak Learning Rate | 1e-5 |
| LR Scheduler | Warmup–Stable–Decay (WSD) |
| Decay Budget | 0 tokens (no decay phase) |
| Final LR Factor (α_f) | 0.0 |
| Optimizer | Decoupled StableAdamW |
| Optimizer β1 / β2 | 0.9 / 0.98 |
| Optimizer ε | 1e-6 |
| Weight Decay | 1e-5 |
| Bias & Norm Weight Decay | Disabled (filter_bias_norm_wd: true) |
| Global Batch Size | 512 sequences |
| Device Microbatch Size (train) | 16 sequences |
| Device Batch Size (eval) | 32 sequences |
| Training Budget | 300,000,000 tokens |
| Sequence Packing | Disabled |
| Padding Strategy | Unpadded (Flash Attention compatible) |
| Count Padding Tokens in Budget | No |
| Batch Size Warmup | From microbatch size up over 30M tokens |
| Attention Dropout | 0.0 (train) / 0.1 (output projection) |
| Seed | 17 |
fill-mask pipeline for masked language modeling, or use it as a backbone for downstream task fine-tuning.pip install transformers torch1from transformers import pipeline
2
3mlm = pipeline("fill-mask", model="proxectonos/MrBERT-nos-gl")
4
5results = mlm("A lingua galega é unha das linguas [MASK] de Europa.")
6for r in results:
7 print(f"{r['token_str']:<20} {r['score']*100:.1f}%")1from transformers import AutoTokenizer, AutoModel
2import torch
3
4tokenizer = AutoTokenizer.from_pretrained("proxectonos/MrBERT-nos-gl")
5model = AutoModel.from_pretrained("proxectonos/MrBERT-nos-gl")
6
7inputs = tokenizer("A lingua galega é unha das linguas romances de Europa.", return_tensors="pt")
8with torch.no_grad():
9 outputs = model(**inputs)
10
11# Mean-pool the last hidden state for a sentence embedding
12embeddings = outputs.last_hidden_state.mean(dim=1)