Qwen3 1.7B Medical Finetuned
This repository contains a bilingual French/English medical LoRA adapter built on top of unsloth/Qwen3-1.7B-unsloth-bnb-4bit.
The training workflow used:
- supervised fine-tuning (SFT) on a curated medical instruction dataset;
- preference alignment with DPO on medical chosen/rejected pairs.
The adapter is intended for experimentation, evaluation, and educational use around medical-domain instruction tuning. It is not a medical device and must not be used as a substitute for a qualified health professional.
Model Details
- Base model:
unsloth/Qwen3-1.7B-unsloth-bnb-4bit
- Adapter type: PEFT LoRA
- Task: causal language modeling / chat-style instruction following
- Languages: French and English
- Final artifact in this folder: DPO-aligned LoRA adapter
- Upstream SFT dataset:
Maphe/medical-sft-5k
- Upstream DPO dataset:
Maphe/medical-dpo-5k
Training setup
The project uses Unsloth, TRL, PEFT, and bitsandbytes with 4-bit loading.
LoRA configuration:
r = 16
lora_alpha = 16
lora_dropout = 0
bias = none
- Target modules:
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
SFT configuration:
- Epochs:
2
- Per-device batch size:
32
- Gradient accumulation:
16
- Learning rate:
2e-4
- Scheduler:
cosine
- Max sequence length:
1024
- Optimizer:
adamw_8bit
- Seed:
42
DPO configuration:
- Epochs:
1
- Per-device batch size:
4
- Gradient accumulation:
8
- Learning rate:
5e-5
- Beta:
0.1
- Scheduler:
cosine
- Max sequence length:
1024
- Optimizer:
adamw_8bit
- Seed:
42
Training Data
Two project datasets were prepared and used in the workflow:
Maphe/medical-sft-5k for supervised fine-tuning
Maphe/medical-dpo-5k for preference optimization
The SFT dataset aggregates bilingual medical QA and MCQ-style examples derived from these Hugging Face sources:
ANR-MALADES/MediQAl
nthngdy/frenchmedmcqa
keivalya/MedQuad-MedicalQnADataset
The DPO dataset is built primarily from:
TsinghuaC3I/UltraMedical-Preference
Project-side preprocessing includes:
- schema normalization across heterogeneous sources;
- prompt/response formatting for chat training;
- deduplication on textual pairs;
- source quota sampling;
- deterministic train/validation/test splitting for SFT;
- heuristic PII anonymization with Presidio and regex-based detectors.
The resulting model is optimized for:
- French and English medical questions;
- short factual answers;
- multiple-choice style medical questions;
- structured, direct responses.
Prompting Format
The training prompt uses a fixed system instruction:
Tu es un assistant medical expert. Reponds de maniere claire, factuelle et structuree. Si la question est en anglais, reponds en anglais.
During training, assistant outputs were formatted in direct-answer mode with an empty Qwen thinking block. This adapter therefore works best with standard chat prompting and concise medical questions.
Intended Uses
Appropriate uses:
- research prototypes in domain adaptation;
- comparison between base and finetuned medical assistants;
- educational work on SFT + DPO pipelines;
- internal experimentation on bilingual medical QA.
Out-of-scope uses:
- diagnosis or treatment decisions without clinician oversight;
- emergency triage;
- autonomous clinical decision support;
- legal, regulatory, or production-grade medical advice systems;
- any workflow requiring guaranteed factuality or safety.
Evaluation
The repository contains a comparative evaluation between the base model and the SFT checkpoint on 500 examples.
Important: the metrics below are for the SFT checkpoint, not for this final DPO adapter. At the time of writing, no dedicated post-DPO benchmark has been added to the repository.
Available evaluation artifacts:
notebooks/eval_results/qwen3_base_vs_sft_output_summary.json
notebooks/eval_results/qwen3_base_vs_sft_output.jsonl
notebooks/eval_results/qwen3_base_vs_sft_output.csv
Summary of SFT-vs-base results:
- Mean METEOR on free-text answers:
0.1361 -> 0.1653 (+0.0292)
- MCQ first-letter score:
0.0515 -> 0.4378 (+0.3863)
- MCQ correct answers:
12 -> 102
Interpretation:
- the finetuning substantially improved MCQ behavior in this project benchmark;
- gains on open-ended generation were positive but more modest;
- automatic metrics remain insufficient to validate clinical quality.
Biases, Risks, and Limitations
This model inherits limitations from both the base model and the medical datasets used during fine-tuning.
Known risks:
- hallucinated or overconfident medical statements;
- incomplete coverage of diseases, populations, and care settings;
- source-data bias toward specific question styles;
- imperfect anonymization in upstream preparation;
- limited evaluation depth;
- possible mismatch between benchmark gains and real clinical usefulness.
This adapter should be used only with strong human review and explicit user-facing warnings.
How to Use
Example with PEFT and Transformers:
1from transformers import AutoModelForCausalLM, AutoTokenizer
2from peft import PeftModel
3
4base_model_id = "unsloth/Qwen3-1.7B-unsloth-bnb-4bit"
5adapter_path = "Maphe/qwen3-1.7b-medical-finetuned"
6
7tokenizer = AutoTokenizer.from_pretrained(base_model_id)
8base_model = AutoModelForCausalLM.from_pretrained(base_model_id)
9model = PeftModel.from_pretrained(base_model, adapter_path)
10
11messages = [
12 {
13 "role": "system",
14 "content": (
15 "Tu es un assistant medical expert. "
16 "Reponds de maniere claire, factuelle et structuree. "
17 "Si la question est en anglais, reponds en anglais."
18 ),
19 },
20 {"role": "user", "content": "Quels sont les symptomes principaux du diabete de type 2 ?"},
21]
22
23prompt = tokenizer.apply_chat_template(
24 messages,
25 tokenize=False,
26 add_generation_prompt=True,
27)
28inputs = tokenizer(prompt, return_tensors="pt")
29outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False)
30print(tokenizer.decode(outputs[0], skip_special_tokens=True))
If you use Unsloth in the same way as in the project notebook, load the base model first and then the LoRA adapter exported in this repository.
Repository Context
This model card is derived from the accompanying project materials:
- root project documentation in
README.md
- training notebook:
notebooks/colab_qwen3_unsloth_finetune.ipynb
- evaluation notebook:
notebooks/colab_qwen3_unsloth_eval_compare.ipynb
The local training artifacts produced by the project include:
- SFT adapter:
notebooks/qwen3-medical-lora/
- DPO adapter:
notebooks/qwen3-medical-dpo-lora/
- SFT checkpoints:
notebooks/sft_output/checkpoint-*
- DPO checkpoint:
notebooks/dpo_output/checkpoint-157
License
No final consolidated license statement has been added yet in the project for the combined derivative artifact. Before public release, verify:
- the license of the base model;
- the license terms of each source dataset;
- whether redistribution of this adapter is compatible with those upstream terms.
Contact
Project owner / publisher: Maphe
If you publish this model publicly, it is worth adding:
- the source repository URL;
- exact dataset revisions;
- a dedicated post-DPO evaluation section;
- explicit medical safety disclaimers in the serving application.
Framework versions