Views
No views yet

1pip install torch transformers regex heapdict
2pip install "git+https://github.com/UniversalComputingResearch/fastboundlessbpe.git@perf/tokenid-training"rustc and cargo; pip builds the extension automatically through Maturin when no compatible wheel is available.1import torch
2from transformers import AutoModelForCausalLM, AutoTokenizer
3
4model_id = "UniversalComputingResearch/Limen0.2B"
5
6tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
7model = AutoModelForCausalLM.from_pretrained(
8 model_id,
9 trust_remote_code=True,
10 torch_dtype=torch.bfloat16,
11).cuda().eval()
12
13prompt = "The future of AI is"
14inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
15with torch.inference_mode():
16 output = model.generate(**inputs, max_new_tokens=64, do_sample=True, temperature=0.8)
17print(tokenizer.decode(output[0], skip_special_tokens=True))trust_remote_code=True is required for the custom model architecture and tokenizer implementation.superword.model is the artifact used during pretraining. Text is encoded without automatic BOS or EOS insertion. Training documents are terminated with <|endoftext|> (ID 16383), which is also the generation stop token. Control-token-like text in ordinary source material is treated as text rather than as markup.theta=100000), RMSNorm, 1,024-token contextbeta1=0.9, beta2=0.95, eps=1e-8). Weight decay was 0.001 and applied only to non-embedding matrix parameters. Gradients were clipped to norm 1.0.0.003, remained constant through 85% of training, and then decayed linearly to zero. The released weights are an FP32 exponential moving average of the final 10% of training (decay=0.999, 4,768 updates), rather than the final raw optimization step.| Source | 0–33.3% | 33.3–66.7% | 66.7–100% |
|---|---|---|---|
| FineWeb-Edu-Dedup | 30% | 30% | 40% |
| Ultra-FineWeb | 50% | 30% | 20% |
| FinePDFs-Edu-English | 5% | 10% | 10% |
| Common-Pile peS2o filtered | 5% | 10% | 10% |
| StackV2 education-filtered | 5% | 10% | 10% |
| Dolma selected non-web | 5% | 10% | 10% |
acc_norm); BLiMP uses pairwise accuracy. Chance-normalized scores map random guessing to 0 and perfect accuracy to 100.| Benchmark | Accuracy | Chance-normalized score |
|---|---|---|
| HellaSwag | 41.98% | 22.64 |
| PIQA | 67.36% | 34.71 |
| ARC-Easy | 53.37% | 37.82 |
| ARC-Challenge | 29.69% | 6.26 |
| CommonsenseQA | 34.15% | 17.69 |
| BLiMP | 83.28% | 66.55 |
acc_norm; BLiMP uses pairwise accuracy. The second line in each row of the accuracy table lists parameter count, vocabulary size, and reported training tokens. Bold indicates the highest benchmark result in each column and the smallest value in each metadata category. Training token counts are reported figures; Qwen2.5's 18T is family-level rather than checkpoint-specific, while GPT-2 Medium's approximately 10B is an estimate.acc_norm)| Model | HellaSwag | PIQA | ARC-Easy | ARC-Challenge | CommonsenseQA | BLiMP |
|---|---|---|---|---|---|---|
| Qwen2.5 0.5B 494M · 151,936 vocab · 18T† | 52.1% | 69.4% | 58.5% | 32.0% | 41.0% | 84.4% |
| Gemma 3 270M 270M · 262,144 vocab · 6T | 41.4% | 68.5% | 57.3% | 28.0% | 42.6% | 82.1% |
| SmolLM2 135M 135M · 49,152 vocab · 2T | 43.2% | 68.2% | 58.7% | 29.7% | 35.5% | 81.4% |
| Limen0.2B 223M · 16,384 vocab · 50B | 42.0% | 67.4% | 53.4% | 29.7% | 34.2% | 83.3% |
| GPT-X2 125M 125M · 32,768 vocab · 75B | 40.5% | 67.1% | 51.6% | 27.6% | 34.6% | 82.9% |
| GPT-2 Medium 355M · 50,257 vocab · ≈10B‡ | 39.3% | 66.6% | 43.5% | 25.0% | 31.7% | 85.2% |
| Model | HellaSwag | PIQA | ARC-Easy | ARC-Challenge | CommonsenseQA | BLiMP |
|---|---|---|---|---|---|---|
| Qwen2.5 0.5B | 36.2 | 38.8 | 44.6 | 9.3 | 26.2 | 68.7 |
| Gemma 3 270M | 21.9 | 37.0 | 43.0 | 4.0 | 28.2 | 64.3 |
| SmolLM2 135M | 24.3 | 36.3 | 44.9 | 6.3 | 19.4 | 62.7 |
| Limen0.2B | 22.6 | 34.7 | 37.8 | 6.3 | 17.7 | 66.6 |
| GPT-X2 125M | 20.7 | 34.2 | 35.5 | 3.4 | 18.3 | 65.8 |
| GPT-2 Medium | 19.1 | 33.2 | 24.6 | 0.0 | 14.6 | 70.4 |
lm-eval, BF16 inference, and batch size 64. Task and macro scores are length-normalized accuracy (acc_norm). Intermediate checkpoints contain raw training weights; the final row reports the EMA model.| Checkpoint | Tokens | Val. loss | BPB | ARC-Easy | ARC-Challenge | HellaSwag | PIQA | Norm. macro |
|---|---|---|---|---|---|---|---|---|
| 1k | 1.049B | 3.0255 | 0.9389 | 38.68% | 22.35% | 29.33% | 59.25% | 37.40% |
| 5k | 5.243B | 2.6211 | 0.8134 | 46.76% | 25.85% | 34.71% | 62.89% | 42.56% |
| 10k | 10.486B | 2.5320 | 0.7857 | 47.47% | 26.71% | 37.39% | 64.85% | 44.11% |
| 20k | 20.972B | 2.4640 | 0.7646 | 51.98% | 27.39% | 38.83% | 65.07% | 45.82% |
| 30k | 31.457B | 2.4313 | 0.7545 | 52.36% | 30.03% | 40.11% | 66.92% | 47.36% |
| 40k | 41.943B | 2.4118 | 0.7484 | 50.51% | 29.44% | 40.76% | 67.03% | 46.93% |
| Final EMA | 50.000B | — | — | 53.37% | 29.69% | 41.98% | 67.36% | 48.10% |
model.safetensors: EMA model weightsconfig.json, config.py, model.py: custom Transformers model definitionsuperword.model, tokenizer_config.json, tokenization_superword.py:
tokenizer model and Transformers integration