Views
No views yet
| Property | Value |
|---|---|
| Architecture | RoBERTa-base |
| Parameters | 124,493,648 |
| Vocabulary | Byte-level BPE, 50,000 tokens |
| Context Window | 512 tokens |
| Hidden Size | 768 |
| Intermediate Size | 3072 |
| Attention Heads | 12 |
| Hidden Layers | 12 |
| Training Data | MaithiliCorpus gold_deduped (47,882 documents, 39.8M words) |
| Training Steps | 14,210 (10 epochs) |
| Effective Batch Size | 32 |
| Optimizer | AdamW (lr=5e-4) |
| Precision | FP16 |
| License | CC-BY-4.0 |
| Property | Value |
|---|---|
| Algorithm | Byte-level BPE |
| Vocabulary Size | 50,000 |
| Training Data | MaithiliCorpus (39.8M words) |
| Training | SentencePiece, --model_type=bpe |
| Special Tokens | <s>, </s>, <unk>, <pad>, <mask> |
| UNK Rate (on Maithili) | 0.00% |
1from transformers import pipeline
2fill = pipeline("fill-mask", model="AiventraLab/maithili-roberta")
3results = fill("भगवान् <mask>क ध्यान")