Views
No views yet
vi)Qwen/Qwen2.5-7B-InstructDIAGNOSIS: Confirmed diagnoses, suspected diagnoses, medical conditions, or diseases.SYMPTOM: Subjective symptoms reported by the patient, or objective clinical signs observed by healthcare providers.TEST_NAME: Name of a laboratory test, imaging study, physical examination, or other clinical examination.TEST_RESULT: Numerical values, categories, or descriptions associated with an examination.DRUG: Medication mention including brand name, generic drug name, dosage/strength, or route.ANATOMY: Anatomical locations, organs, or body structures.PROCEDURE: Medical procedures, surgeries, therapies, or interventions.DEMOGRAPHICS: Patient-specific identifiers such as age, gender, occupation, or location.is_negated: true only if the entity is explicitly denied (e.g., "không", "chưa", "phủ nhận", "loại trừ").is_uncertain: true if the entity is suspected, possible, or hedged rather than confirmed (e.g., "nghi ngờ", "có thể", "theo dõi", "chưa loại trừ").severity: A category describing intensity when stated (e.g., "nhẹ", "vừa", "nặng"); null if not mentioned.temporality: Whether the entity is "current" (present/ongoing) or "recent" (a recent but resolved occurrence).is_historical: true if the entity refers to a past occurrence (e.g., "cách đây", "tiền sử", "đã từng").is_family: true only if the entity belongs to a family member (e.g., "bố", "mẹ", "gia đình") rather than the patient.is_conditional: true if the entity is hypothetical or conditional rather than an actual finding (e.g., "nếu", "khi", "trong trường hợp").system prompt, and the model will strictly adhere to extracting only those requested types into the required JSON schema.is_negated, is_uncertain, severity, temporality, is_historical, is_family, is_conditional) rely on the model's interpretation of surrounding language and may be misassigned in complex, nested, or ambiguous clinical phrasing.1import json
2from transformers import AutoModelForCausalLM, AutoTokenizer
3
4model_id = "PeterPaker123/Qwen2.5-7B-Vietnamese-Medical-NER"
5
6tokenizer = AutoTokenizer.from_pretrained(model_id)
7model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto", torch_dtype="auto")
8
9system_prompt = """You are a medical expert. Your task is to identify and extract entities from the given text.
10The entity types to extract are:
11- DIAGNOSIS: confirmed diagnoses, suspected diagnoses, medical conditions, or diseases.
12- SYMPTOM: subjective symptoms reported by the patient (CƠ_NĂNG) or objective clinical signs observed by healthcare providers (THỰC_THỂ).
13- TEST_NAME: name of a laboratory test, imaging study, physical examination, or other clinical examination.
14- TEST_RESULT: numerical values, categories, or descriptions associated with an examination.
15- DRUG: medication mention including brand name, generic drug name, dosage/strength, route.
16- ANATOMY: anatomical locations, organs, or body structures.
17- PROCEDURE: medical procedures, surgeries, therapies, or interventions.
18- DEMOGRAPHICS: patient-specific identifiers such as age, gender, occupation, or location.
19
20For each extracted entity, evaluate the text context and determine the following boolean flags:
211. "is_negated": true ONLY IF explicitly denied ("không", "chưa", "phủ nhận", "loại trừ"). Otherwise, false.
222. "is_uncertain": true if suspected, possible, or hedged rather than confirmed ("nghi ngờ", "có thể", "theo dõi", "chưa loại trừ"). Otherwise, false.
233. "severity": a category describing intensity when stated ("nhẹ", "vừa", "nặng"). null if not mentioned.
244. "temporality": "current" if present/ongoing, or "recent" if a recent but resolved occurrence.
255. "is_historical": true if it occurred in the past ("cách đây", "tiền sử", "đã từng"). Otherwise, false.
266. "is_family": true ONLY IF it belongs to a family member ("bố", "mẹ", "gia đình"). False if it belongs to the patient.
277. "is_conditional": true if hypothetical or conditional rather than an actual finding ("nếu", "khi", "trong trường hợp"). Otherwise, false.
28
29Return the result as a JSON list containing dictionaries with "entity", "type", "is_negated", "is_uncertain", "severity", "temporality", "is_historical", "is_family", and "is_conditional" keys. If no relevant entities are found, return an empty list []."""
30
31prompt = [
32 {
33 "role": "system",
34 "content": system_prompt
35 },
36 {
37 "role": "user",
38 "content": "Text: Bệnh nhân nam 45 tuổi vào viện vì đau tức ngực trái dữ dội. Tiền sử tăng huyết áp, đang dùng Amlodipine 5mg. Khám phổi không ran. Theo dõi nhồi máu cơ tim. Chỉ định chụp mạch vành. Bố bệnh nhân bị đái tháo đường."
39 }
40]
41
42inputs = tokenizer.apply_chat_template(prompt, return_tensors="pt", add_generation_prompt=True).to(model.device)
43outputs = model.generate(inputs, max_new_tokens=768, temperature=0.0)
44
45response = tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True)
46print(response)
47# Expected Output:
48# [
49# {"entity": "nam", "type": "DEMOGRAPHICS", "is_negated": false, "is_uncertain": false, "severity": null, "temporality": "current", "is_historical": false, "is_family": false, "is_conditional": false},
50# {"entity": "45 tuổi", "type": "DEMOGRAPHICS", "is_negated": false, "is_uncertain": false, "severity": null, "temporality": "current", "is_historical": false, "is_family": false, "is_conditional": false},
51# {"entity": "đau tức ngực trái dữ dội", "type": "SYMPTOM", "is_negated": false, "is_uncertain": false, "severity": "nặng", "temporality": "current", "is_historical": false, "is_family": false, "is_conditional": false},
52# {"entity": "ngực trái", "type": "ANATOMY", "is_negated": false, "is_uncertain": false, "severity": null, "temporality": "current", "is_historical": false, "is_family": false, "is_conditional": false},
53# {"entity": "tăng huyết áp", "type": "DIAGNOSIS", "is_negated": false, "is_uncertain": false, "severity": null, "temporality": "current", "is_historical": true, "is_family": false, "is_conditional": false},
54# {"entity": "Amlodipine 5mg", "type": "DRUG", "is_negated": false, "is_uncertain": false, "severity": null, "temporality": "current", "is_historical": false, "is_family": false, "is_conditional": false},
55# {"entity": "ran", "type": "SYMPTOM", "is_negated": true, "is_uncertain": false, "severity": null, "temporality": "current", "is_historical": false, "is_family": false, "is_conditional": false},
56# {"entity": "nhồi máu cơ tim", "type": "DIAGNOSIS", "is_negated": false, "is_uncertain": true, "severity": null, "temporality": "current", "is_historical": false, "is_family": false, "is_conditional": false},
57# {"entity": "chụp mạch vành", "type": "PROCEDURE", "is_negated": false, "is_uncertain": false, "severity": null, "temporality": "current", "is_historical": false, "is_family": false, "is_conditional": false},
58# {"entity": "đái tháo đường", "type": "DIAGNOSIS", "is_negated": false, "is_uncertain": false, "severity": null, "temporality": "current", "is_historical": false, "is_family": true, "is_conditional": false}
59# ]| Dataset Name | Domain / Source | Total Samples | Entities Used During SFT |
|---|---|---|---|
| VietBioNER | Tuberculosis (TB) clinical treatment | ~1,700 | Symptom & Disease, Diagnostic Procedure, Location, Date |
| PhoNER_COVID19 | COVID-19 pandemic surveillance | ~10,000 | Symptom & Disease, Patient ID, Age, Gender, Location, Org |
| ViMQ | Healthcare dialogue systems | ~10,000 | Symptom, Disease, Drug, Procedure, Body Structure |
| ViMedNER | General Medical NER | ~10,000 | Symptom, Disease, Drug, Body Part, Procedure, Time |
bf16 mixed precisiontransformerstrl (SFTTrainer)torch (with sdpa attention)1@inproceedings{vietbioner,
2 title = "{A Named Entity Recognition Corpus for Vietnamese Biomedical Texts to Support Tuberculosis Treatment}",
3 author = "Phan, Uyen and Nguyen, Phuong and Nguyen, Nhung",
4 booktitle = "Proceedings of the 13th Language Resources and Evaluation Conference",
5 year = "2022",
6 publisher = "European Language Resources Association"
7}
8
9@inproceedings{PhoNER_COVID19,
10 title = {{COVID-19 Named Entity Recognition for Vietnamese}},
11 author = {Thinh Hung Truong and Mai Hoang Dao and Dat Quoc Nguyen},
12 booktitle = {Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies},
13 year = {2021}
14}
15
16@article{vimq2023,
17 title={ViMQ: A Vietnamese Medical Question Dataset for Healthcare Dialogue System Development},
18 author={Vu, Hieu M. and Phan, Long and Nguyen, Nhung and others},
19 year={2023}
20}
21
22@article{10.4108/eetinis.v11i3.5221,
23 author={Pham Van Duong and Tien-Dat Trinh and Minh-Tien Nguyen and Huy-The Vu and Minh Chuan Pham and Tran Manh Tuan and Le Hoang Son},
24 title={ViMedNER: A Medical Named Entity Recognition Dataset for Vietnamese},
25 journal={EAI Endorsed Transactions on Industrial Networks and Intelligent Systems},
26 volume={11},
27 number={4},
28 publisher={EAI},
29 year={2024},
30 doi={10.4108/eetinis.v11i3.5221}
31}is_negated, is_uncertain, severity, temporality, is_historical, is_family, is_conditional) that qualify each extracted entity based on its clinical context.