OTel-LLM-12B-IT
OTel-LLM-12B-IT is a context-grounded telecom language model full-parameter fine-tuned on OTel telecommunications data. It is part of the
OTel Family of Models, an open-source initiative to build reference AI resources for the global telecommunications sector.
Across the core OTel LLM baselines, OTel fine-tuning improves context-grounded correctness over the base checkpoints by +3.7 to +10.0 percentage points.
Community Use
As of June 23, 2026, the released OTel models had more than 18 million downloads, and the Open Telco AI project had received 157+ pieces of media coverage worldwide.
Model Details
| Attribute | Value |
|---|
| Base model | google/gemma-3-12b-it |
| Parameters | 12B |
| OTel training dataset | OTel-LLM |
| Dataset fields | prompt, completion, abstention, chunk-count metadata, token-count metadata |
| Training method | Full-parameter post-training / fine-tuning |
| Language | English |
| OTel release license | Apache 2.0 |
Model Lineage
google/gemma-3-12b-it -> OTel-LLM full-parameter post-training -> farbodtavakkoli/OTel-LLM-12B-IT
OTel vs. Base Model
| Metric | Base model | OTel fine-tuned | Delta | Evaluation split |
|---|
| LLM-as-judge correctness | 78.3% | 83.3% +/- 0.5 | +5.0 pp | OTel-LLM held-out 10% |
Standard errors are computed with bootstrap resampling (n=10) over the held-out OTel evaluation partition. LLM correctness is judged by GPT-4o mini using the retrieved context and reference answer.
Evaluation Caveats
- LLM results measure context-grounded answer generation from retrieved context, not unrestricted context-free telecom QA.
- Reported standard errors come from bootstrap resampling over the held-out evaluation partitions.
- Answer quality depends on the retriever, reranker, context window, and prompt policy around the model.
- External benchmark transfer, multilingual performance, and per-subdomain performance should be evaluated separately for production settings.
Training Data
The model was trained on telecom-focused data curated by 100+ domain experts. The raw corpus contained roughly 1.1M training points and was filtered to 326,767 higher-confidence examples.
| Source | Contributor |
|---|
| arXiv telecom papers, 3GPP standards, telecom Wikipedia, telecom Common Crawl | Yale University |
| GSMA Permanent Reference Documents, Discover portal | GSMA |
| IETF RFC series | NetoAI |
| Industry whitepapers | Khalifa University |
| O-RAN specifications (working groups 1, 2, 4, 5, 6, 7, 8, 9, 10) | University of Leeds |
| O-RAN documents across working groups | The University of Texas at Dallas |
The OTel datasets release derived QA/retrieval/reranking examples rather than the raw source documents.
Each released dataset includes a dataset card and Croissant metadata with Responsible AI fields for data limitations, biases, sensitive-information considerations, use cases, social impact, synthetic-data status, and provenance.
Representative Training Row
OTel-LLM rows pair a context-grounded telecom RAG prompt with a reference completion.
1{
2 "anchor": "How can a cell be considered to be operating in MBSFN mode for 3.84/7.68 Mcps TDD?",
3 "completion": "A cell shall be considered to be operating in MBSFN mode when individual scrambling codes are assigned to all timeslots via the IE \"TDD MBSFN Information\".",
4 "abstention": false,
5 "n_positive_chunks": 1,
6 "n_negative_chunks": 4
7}
Intended Use
This model is intended for context-grounded telecom answer generation in Retrieval-Augmented Generation (RAG) pipelines. It should receive retrieved telecom context and generate an answer grounded in that context.
The model is not optimized for unrestricted context-free question answering. For questions where the retrieved context is missing or insufficient, use an abstention-aware prompt or one of the dedicated -Safety variants.
Training Recipe
| Item | Value |
|---|
| Framework | ScalarLM |
| Optimizer | AdamW, 8-bit |
| Learning-rate schedule | Cosine decay with warmup |
| Weight decay | 0.01 |
| Warmup steps | 100 |
| Random seed | 42 |
| Maximum sequence length | 1500 tokens |
| Precision | BF16 |
| Attention | Flash Attention 2 |
| Distributed training | Fully Sharded Data Parallel |
| Gradient checkpointing | Enabled |
| Epochs | 3 for LLM/embedding models; 2 for rerankers |
| Compute | AMD MI300X/MI325X/MI355X and NVIDIA A100/H100 GPUs |
Usage
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4model_name = "farbodtavakkoli/OTel-LLM-12B-IT"
5tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
6model = AutoModelForCausalLM.from_pretrained(
7 model_name,
8 torch_dtype=torch.bfloat16,
9 device_map="auto",
10 trust_remote_code=True,
11)
12
13prompt = """You are a precise telecom assistant in a RAG pipeline.
14Use only the retrieved context to answer.
15
16User Question
17What is the purpose of the F1 interface in O-RAN?
18
19Retrieved Contexts
20CONTEXT 1
21The F1 interface connects the O-RAN Distributed Unit (O-DU) to the O-RAN Central Unit (O-CU).
22
23Answer:"""
24
25inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
26outputs = model.generate(**inputs, max_new_tokens=256)
27print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Limitations and Responsible Use
- OTel models are domain-specific to telecommunications and should not be treated as general-purpose models.
- The current release is English-only and primarily text-centric.
- The reported OTel performance results use held-out OTel evaluation partitions and should not be interpreted as results from a fully independent external benchmark suite.
- Aggregate scores can hide subdomain variation; collaborator stress tests suggest O-RAN retrieval is comparatively strong, while academic-paper and GSMA PRD examples need further curation.
- Generated telecom content should be verified before operational, customer-facing, regulatory, safety, or network-configuration use.
- Users must comply with both the OTel release license and the upstream base-model license or terms.
- For unrestricted telecom QA without retrieved context, use a separately evaluated context-free QnA model rather than assuming this RAG-oriented checkpoint will behave optimally.
Related Models
Project Resources
Citation
1@misc{otel_models_2026,
2 title = {OTel: Open Telco AI Datasets, Benchmarks, and Models},
3 author = {Tavakkoli, Farbod and others},
4 year = {2026},
5 note = {Open Telco (OTel) model release},
6 url = {https://huggingface.co/farbodtavakkoli}
7}
Contact