Views
No views yet
| Field | Value |
|---|---|
| Base model | google/gemma-3-270m-it |
| Training method | Full-FT (CE loss) |
| Training data | DIA-GUARD splits (~836K train, 178K val) |
| Domain | LLM safety classification across 48 English dialects |
| Role | Student model (used as KD student in DIA-GUARD pipeline) |
| License | Gemma Terms of Use (inherited from base model) |
safe or unsafe across English dialects1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model = AutoModelForCausalLM.from_pretrained("jsl5710/Shield-Gemma-3-270m-Full-FT-CE", torch_dtype="bfloat16")
4tokenizer = AutoTokenizer.from_pretrained("jsl5710/Shield-Gemma-3-270m-Full-FT-CE")
5
6prompt = "<your prompt here>"
7inputs = tokenizer.apply_chat_template(
8 [{"role": "system", "content": "You are DIA-Guard, a multilingual safety assistant."},
9 {"role": "user", "content": prompt}],
10 return_tensors="pt", add_generation_prompt=True,
11)
12outputs = model.generate(inputs, max_new_tokens=4)
13print(tokenizer.decode(outputs[0], skip_special_tokens=True))
14# Expected: 'safe' or 'unsafe'| Metric | Value |
|---|---|
| Final epoch | 0.73/3 (early-stopped) |
| Train loss | 0.5839 |
| Train accuracy | 87.29% |
| Eval loss | 1.078 |
| Eval accuracy | 79.68% |
| Batch size (per_device × grad_accum) | 256 × 1 = 256 |
| Liger Kernel | ✅ enabled |
| Stopped via | EarlyStoppingCallback (patience=3, metric=eval_loss) |
Eval was performed on a 2,000-sample subset of the DIA-GUARD val split (full val: 178K samples). Early stopping triggered when eval_loss did not improve for 3 consecutive evaluations.
| Metric | Value |
|---|---|
| Test Accuracy | 0.9654 |
| Macro Precision | 0.9676 |
| Macro Recall | 0.9634 |
| Macro F1 | 0.9650 |
| Support | 181,874 |
| Class | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| safe | 0.9844 | 0.9392 | 0.9613 | 83,140 |
| unsafe | 0.9507 | 0.9875 | 0.9688 | 98,734 |
| Pred safe | Pred unsafe | |
|---|---|---|
| True safe | 78,087 | 5,053 |
| True unsafe | 1,234 | 97,500 |
Per-dialect breakdown available inper_dialect.jsonin the corresponding results folder.
jsl5710/Shield1@misc{diaguard2026,
2 title = {DIA-GUARD: Dialect-Informed Adversarial Guard for LLM Safety},
3 author = {Jason Lucas et al.},
4 year = {2026},
5 howpublished = {\url{https://github.com/jsl5710/dia-guard}}
6}