PatientAgent: patientagent-dpo-sft128
Dpo adapter trained over the rank-128 sft model for research on simulated patient dialogue. The adapter is based on Qwen/Qwen3.5-4B and was trained on MTS-Dialog-derived data.
Composition and loading
This repository stores a PEFT LoRA adapter, not a standalone merged model. Its adapter configuration reports LoRA rank 16 and alpha 32.
Data and evaluation
The current published G-Eval run evaluates 174 eligible dialogues from MTS-Dialog test1, excluding source dialogues with no patient turn. Consult the evaluation dataset for exact scores and score directions.
Limitations and intended use
Research benchmark only. This model is not a medical device and must not be used as medical advice or as evidence that another model is clinically safe. It can hallucinate, omit facts, or produce culturally and linguistically narrow dialogue. It was developed for English-language research data and has not been clinically validated.
MTS-Dialog contains professionally written simulated dialogues rather than recordings of real patients. Users remain responsible for complying with the source dataset and Qwen licenses.