Med-LLaMA3.1-8B — Medical QLoRA Adapter (LoRA weights only)
Parameter-efficient medical adaptation of Llama-3.1-8B using QLoRA (4-bit NF4 + LoRA).
This repository contains the LoRA adapter only — it must be applied on top of the base
model at load time. For a ready-to-use, standalone checkpoint, see the merged version
linked below.
This is the 8B (high-capacity flagship) member of the Med-LLaMA3 family introduced in the paper
“Med-LLaMA3: Advancing Medical Question-Answering Through Parameter-Efficient Fine-Tuning of Large
Language Models” (Applied Sciences, 2026). The family adapts the LLaMA-3 architecture to the medical
domain by training only a small fraction of the base model’s parameters (4.01% for this 8B variant),
achieving strong medical question-answering performance while reducing memory use by roughly 75% via
4-bit quantization — enabling development and inference on low-cost, consumer-grade hardware.
The 8B variant is the high-capacity model for complex clinical reasoning. It attains a mean
accuracy of 75.71% across the eight MMLU medical subsets and performs comparably to the
institutionally trained LLaMA3-Med42-8B at the same scale, while being trained on consumer-grade GPUs.
ℹ️ Base checkpoint. A LoRA adapter only loads correctly onto the exact base model it was trained
on. This adapter targets the instruct checkpoint meta-llama/Llama-3.1-8B-Instruct (note: the 8B
uses LLaMA 3.1, not 3.2). Use that same base in the code below.
Intended uses
Primary use cases
Medical question answering (multiple-choice and open-ended) — the highest-accuracy variant in the family.
Clinical knowledge lookup and clinical decision support assistance.
(A pre-merged checkpoint is also published separately — see the link at the top of this card.)
Training data
The Med-LLaMA3 family was fine-tuned on a curated medical instruction dataset of over 1.5 million
samples, organized along a three-axis taxonomy: source type (examination QA, clinical dialogue,
biomedical literature, encyclopedic reference) × clinical granularity (basic science, clinical
reasoning, patient communication) × task format (multiple-choice, open-ended QA, generative
dialogue). All sources were consolidated into a unified instruction–response schema
(system, context, question, answer, choices).
Sources include:
MedAlpaca / Medical Meadow collection — MEDIQA, Medical Flashcards, WikiDoc, WikiDoc Patient
Information, MedQA, CORD-19, and PubMed Causal subsets
MedMCQA — Indian medical entrance exam (AIIMS & NEET PG) multiple-choice questions
Evaluation integrity: The eight MMLU medical subsets were used only for held-out
evaluation and were excluded from the fine-tuning corpus. For benchmarks with official splits
(MedMCQA, MedQA-USMLE, PubMedQA), only the official training partitions were used for fine-tuning.
Training procedure
LoRA and optimization settings are identical across the 1B, 3B, and 8B variants; sequence length, batch
size, and gradient accumulation are scaled to each model’s memory footprint. The settings below are for
the 8B variant.
Setting
Value (8B)
Method
QLoRA (4-bit NF4 base, LoRA adapters in higher precision)
float16 · gradient checkpointing (no DeepSpeed / FlashAttention — unsupported on T4)
Hardware
2 × NVIDIA T4 (30 GB total, free-tier Kaggle); ~15–25 days per run
Experiment tracking
Weights & Biases
The QLoRA recipe keeps the base weights frozen and quantized, allocating optimizer state only for the
335 M LoRA parameters rather than the 8 B base — which is what makes fine-tuning feasible on a 30 GB
2× T4 setup.
Evaluation
All re-run models (including the baselines) were evaluated with the EleutherAI LM Evaluation Harness
under identical conditions: same harness version (v0.4.2), identical prompt templates, and 5-shot
prompting. Reported ± values are 95% bootstrap confidence intervals (1000 resamples); statistical
significance uses McNemar’s test on per-item paired correctness.
MMLU medical subsets (5-shot accuracy %)
Comparison against the institutionally trained LLaMA3-Med42-8B under identical conditions
(this is the comparison visualized in Figure 3 of the paper). Family context: mean accuracy scales with
size — 1B = 48.64%, 3B = 64.24%, 8B = 75.71%.
MMLU medical subset
Med-LLaMA3.1-8B (Ours)
LLaMA3-Med42-8B
Anatomy
71.11 (±3.92)
69.63 (±3.97)
Clinical Knowledge
79.62 (±2.48)
76.60 (±2.61)
College Biology
84.03 (±3.06)
81.25 (±3.26)
College Medicine
67.63 (±3.57)
67.05 (±3.58)
Medical Genetics
84.00 (±3.68)
76.00 (±4.29)
Nutrition
83.66 (±2.12)
72.88 (±2.55)
Professional Medicine
77.21 (±2.55)
75.00 (±2.63)
Virology
58.43 (±3.84)
49.40 (±3.89)
Mean (8 subsets)
75.71
71.10
Statistical interpretation (honest framing):
vs. Llama-3.1-8B-Instruct (untuned base): improvements are statistically significant on Anatomy
(p=0.021), Clinical Knowledge (p=0.003), Medical Genetics (p<0.001), Nutrition (p<0.001), and
Professional Medicine (p=0.008); College Biology (p=0.11) and College Medicine (p=0.43) are not
significant. After Bonferroni correction across the eight subsets (per-subset threshold 0.00625),
Clinical Knowledge, Medical Genetics, and Nutrition remain significant; Anatomy and Professional
Medicine are significant only before correction.
vs. LLaMA3-Med42-8B (institutionally trained): differences are not statistically significant
— i.e., comparable performance, not demonstrated superiority, at the same parameter scale.
Medical QA benchmarks (5-shot accuracy %)
Model
MedMCQA
MedQA
PubMedQA
Med-LLaMA3.1-8B (Ours)
61.8 (±0.75)
62.4 (±1.37)
77.0 (±1.96)
LLaMA3-Med42-8B
60.3 (±0.76)
62.8 (±1.36)
77.0 (±1.88)
Llama-3.1-8B-Instruct
59.6 (±0.76)
61.9 (±1.36)
75.8 (±1.92)
MedGemma-4B-it
32.2 (±0.72)
27.7 (±1.26)
55.2 (±2.23)
The +2.2-point gain on MedMCQA over Llama-3.1-8B-Instruct is statistically significant (p=0.004).
The difference vs. LLaMA3-Med42-8B on MedMCQA is not significant (p=0.38). On MedQA and PubMedQA the
models are statistically tied.
Approximately 3.58 GB model size and ~5.2 GB GPU memory allocated at inference, ~30 tokens/s,
with ~20 ms first-token latency — comparable to other 8B 4-bit baselines and well within consumer-GPU
budgets.
See the paper for full tables, generation-quality metrics,
expert evaluation, and the safety analysis.
Limitations & responsible use
Not a medical device. This model is a research artifact. It must not be used for autonomous
diagnosis, treatment, prescribing, or any decision affecting patient care without review by a
qualified healthcare professional. In expert review of 100 generative cases, 94% were rated “Good” but
6% contained critical errors — underscoring the need for human verification.
Comparable, not superior. Against the institutionally trained LLaMA3-Med42-8B, differences are
not statistically significant; gains over the untuned base are significant on some subsets but not all
(see above). On MedQA and PubMedQA there is no significant advantage.
Hallucination & over-elaboration. The model can produce fluent but incorrect information and tends
to elaborate beyond the question, which may obscure key points in time-sensitive settings.
Abbreviation ambiguity. A known high-severity error source. The paper’s safety pilot shows that
context-disambiguation preprocessing reduces abbreviation-ambiguity errors from 30% to 10% on a
held-out set; consider applying similar preprocessing.
Data & bias. Training data may under-represent certain populations, conditions, or regional
practices, and may encode biases present in the source corpora. Rare-condition coverage is limited.
Privacy & compliance. Do not input protected health information (PHI) unless your deployment is
appropriately secured and compliant with applicable regulations (e.g., HIPAA, GDPR).
Evaluation scope. Benchmarks are English-only and dominated by multiple-choice formats; long-form
reasoning, multi-turn dialogue safety, non-English text, and out-of-distribution robustness are not
evaluated.
Recommended deployment: clinical decision support (not autonomous decisions), educational tool,
documentation assistant (with physician review), and literature synthesis.
License
This adapter is released under the Llama 3.1 Community License,
inherited from the base model. By using it you agree to Meta’s Llama 3.1 license terms and
Acceptable Use Policy. Review the licenses of the individual training datasets for any additional
restrictions on derived use.
Citation
If you use this model, please cite the paper:
bibtex
1@article{aboelenen2026medllama3,
2 title = {Med-LLaMA3: Advancing Medical Question-Answering Through Parameter-Efficient Fine-Tuning of Large Language Models},
3 author = {Abo El-Enen, Mohamed Ahmed and Ismail, Sally S. and Nazmy, Taymoor Mohamed},
4 journal = {Applied Sciences},
5 volume = {16},
6 number = {12},
7 pages = {6158},
8 year = {2026},
9 publisher = {MDPI},
10 doi = {10.3390/app16126158},
11 url = {https://www.mdpi.com/2076-3417/16/12/6158}
12}
Authors & contact
Mohamed Ahmed Abo El-Enen, Sally S. Ismail, and Taymoor Mohamed Nazmy
Faculty of Computer and Information Sciences, Ain Shams University, Cairo, Egypt.