moe-expert-research-4b is a specialized 4-billion parameter Small Language Model (SLM), LoRA fine-tuned on the LUMI-G Supercomputer (8× AMD Instinct™ MI250X GCDs (4× physical modules, 64GB HBM2e per GCD)).
Within the MoE Sovereign compound AI architecture, this model serves as the Evidence Synthesis, Literature Analysis, and Citation Verification Expert. It is trained to perform comparative document analysis, extract empirical metrics from research benchmarks, evaluate engineering trade-offs, and produce structured analytical syntheses where every substantive claim is tightly bounded to retrieved context spans.
🎯 Functional Scope & Capabilities
Evidence-Grounded Synthesis: Compiles multi-document source texts into cohesive comparative analyses without introducing ungrounded external claims.
Constrained Citation Generation: Formulates references and factual attributions anchored exclusively to provided source spans ([Doc ID: Section]).
Engineering Trade-Off Evaluation: Structures multi-dimensional trade-off matrices (Latency vs. Throughput, Consistency vs. Availability, Memory vs. Compute).
State-of-the-Art Surveying: Summarizes architectural evolutions across computer science, distributed systems, and AI systems research.
🎯 Training Objectives & Intended Behavioral Specialization
Capability
Base Stock Qwen 3.5 4B
moe-expert-research-4b (Distilled)
Citation Grounding
Invents non-existent DOIs, authors, or paper titles
Strict In-Context Provenance; cites exact document tags and section anchors
Hierarchical Technical Architecture with clear structural transitions
📊 Empirical Evaluation (Held-Out Benchmark Suite)
ℹ️ Evaluation Status: Evaluated on held-out validation splits ($N=1,000$, zero training contamination). Full cross-architecture ablation suites across Compound AI vs. Monolithic LLMs are undergoing active execution in the Sovereign Scientific Benchmark Suite v1.
Evaluated on a held-out benchmark suite of 1,000 multi-document research synthesis tasks with zero training contamination, evaluated across factual entailment (NLI) and citation verification:
Evaluation Metric
Base Stock Qwen 3.5 4B
moe-expert-research-4b (Distilled)
Delta ($\Delta$)
Citation Precision (Valid Provenance)
62.1 %
96.8 %
+34.7 %
Claim-Evidence Entailment (NLI Hold)
69.4 %
95.1 %
+25.7 %
Hallucinated Fact Ratio
16.8 %
2.4 %
-14.4 %
Multi-Source Trade-Off Completeness
58.0 %
91.4 %
+33.4 %
Structured Matrix Formatting Fidelity
74.2 %
98.0 %
+23.8 %
Long-Context Context Span Retrieval
63.5 %
93.2 %
+29.7 %
Note: Evaluated at temperature=0.15 across 3 independent seeds. Citation precision measures the percentage of generated citations that accurately point to supporting evidence in the source text.
Closed-World Retrieval Constraint: When zero retrieval context is provided, the model explicitly acknowledges lack of evidence rather than generating probabilistic general-knowledge claims.
Conflicting Primary Sources: When input documents present mutually contradictory empirical findings, the model highlights the contradiction for the Judge oracle rather than attempting unilateral arbitration.
Context Length Budgeting: For document corpora exceeding 64k tokens, iterative chunking via the MoE Sovereign compound pipeline is recommended for maximum extraction recall.
💻 Quickstart Guide (Ollama & Llama.cpp)
1. Ollama Modelfile
dockerfile
1FROM ./moe-expert-research-4b-Q4_K_M.gguf2PARAMETER num_ctx 262144
3PARAMETER temperature 0.15
4TEMPLATE """{{ if .System }}<|im_start|>system
5{{ .System }}<|im_end|>
6{{ end }}{{ if .Prompt }}<|im_start|>user
7{{ .Prompt }}<|im_end|>
8{{ end }}<|im_start|>assistant
9{{ .Response }}<|im_end|>"""
2. Python Inference
python
1import torch
2from transformers import AutoModelForCausalLM, AutoTokenizer
34model_id ="h3rb3rn/moe-expert-research-4b"56tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)7model = AutoModelForCausalLM.from_pretrained(8 model_id,9 torch_dtype=torch.bfloat16,10 device_map="auto",11 trust_remote_code=True12)1314prompt ="<|im_start|>user\nSynthesize the technical trade-offs between Raft and Paxos based on the provided papers, with exact citation tags.<|im_end|>\n<|im_start|>assistant\n"15inputs = tokenizer(prompt, return_tensors="pt").to(model.device)16outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.15)17print(tokenizer.decode(outputs[0], skip_special_tokens=True))
📑 Citation
bibtex
1@misc{moe_sovereign_2026_research4b,
2 author = {Horn, Philipp and MoE Sovereign Core AI Team},
3 title = {MoE Sovereign Research Expert 4B: Evidence-Grounded Synthesis SLM},
4 year = {2026},
5 publisher = {Hugging Face},
6 howpublished = {\url{https://huggingface.co/h3rb3rn/moe-expert-research-4b}},
7 note = {Trained on the EuroHPC LUMI-G Supercomputer}
8}