Views
No views yet
| Field | Value |
|---|---|
| Base model | Qwen/Qwen3Guard-Gen-0.6B |
| Training method | Full-FT (CE loss) |
| Training data | DIA-GUARD splits (~836K train, 178K val) |
| Domain | LLM safety classification across 48 English dialects |
| Role | Student model (used as KD student in DIA-GUARD pipeline) |
| License | Apache 2.0 (inherited from base model) |
safe or unsafe across English dialects1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model = AutoModelForCausalLM.from_pretrained("jsl5710/Shield-Qwen3Guard-Gen-0.6B-Full-FT-CE", torch_dtype="bfloat16")
4tokenizer = AutoTokenizer.from_pretrained("jsl5710/Shield-Qwen3Guard-Gen-0.6B-Full-FT-CE")
5
6prompt = "<your prompt here>"
7inputs = tokenizer.apply_chat_template(
8 [{"role": "system", "content": "You are DIA-Guard, a multilingual safety assistant."},
9 {"role": "user", "content": prompt}],
10 return_tensors="pt", add_generation_prompt=True,
11)
12outputs = model.generate(inputs, max_new_tokens=4)
13print(tokenizer.decode(outputs[0], skip_special_tokens=True))
14# Expected: 'safe' or 'unsafe'| Metric | Value |
|---|---|
| Final epoch | 0.52/3 (early-stopped) |
| Train loss | 0.0977 |
| Train accuracy | 97.9% |
| Eval loss | 0.17 |
| Eval accuracy | 96.8% |
| Batch size (per_device × grad_accum) | 128 × 1 = 128 |
| Liger Kernel | ✅ enabled |
| Stopped via | EarlyStoppingCallback (patience=3, metric=eval_loss) |
Eval was performed on a 2,000-sample subset of the DIA-GUARD val split (full val: 178K samples). Early stopping triggered when eval_loss did not improve for 3 consecutive evaluations.
| Metric | Value |
|---|---|
| Test Accuracy | 0.5432 |
| Macro Precision | 0.5573 |
| Macro Recall | 0.5005 |
| Macro F1 | 0.3545 |
| Support | 181,874 |
| Class | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| safe | 0.5714 | 0.0027 | 0.0055 | 83,140 |
| unsafe | 0.5431 | 0.9983 | 0.7035 | 98,734 |
| Pred safe | Pred unsafe | |
|---|---|---|
| True safe | 228 | 82,912 |
| True unsafe | 171 | 98,563 |
Per-dialect breakdown available inper_dialect.jsonin the corresponding results folder.
jsl5710/Shield1@misc{diaguard2026,
2 title = {DIA-GUARD: Dialect-Informed Adversarial Guard for LLM Safety},
3 author = {Jason Lucas et al.},
4 year = {2026},
5 howpublished = {\url{https://github.com/jsl5710/dia-guard}}
6}