Fixes typos, missed accents, and OCR-style char errors in Vietnamese
text in one pass: Toi yu Vit Nam → Tôi yêu Việt Nam. Strictly more
than diacritic restoration — handles letter-level mistakes, missing /
extra characters, and OCR substitutions like o↔0, l↔1, m↔rn.
Fine-tuned from
vinai/bartpho-syllable-base on the
nrl-ai/vn-spell-correction-train
corpus (459K (noisy, clean) Vietnamese pairs synthesized from a
register-balanced Wiki+news mix via
nom.text.noise).
1from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
2
3tok = AutoTokenizer.from_pretrained("nrl-ai/vn-spell-correction-small")
4model = AutoModelForSeq2SeqLM.from_pretrained("nrl-ai/vn-spell-correction-small").eval()
5
6text = "Toi yu Vit Nam"
7out = model.generate(**tok(text, return_tensors="pt"), max_length=256)
8print(tok.decode(out[0], skip_special_tokens=True))
9# Tôi yêu Việt Nam
1from nom.text.diacritic_models import HFDiacriticModel
2restorer = HFDiacriticModel(model_id="nrl-ai/vn-spell-correction-small")
3fixed = restorer.predict_batch(noisy_sentences, batch_size=16)
Evaluation uses
nrl-ai/vn-spell-correction-eval
(2,098 pairs across 4 registers x 2 noise levels). Word accuracy after
NFC + punctuation normalization on both sides.
Where this model sits in the public Vietnamese spell-correction
landscape — same 8-split grid for every measured row.
-
In-distribution metric, real-world is harder — measured. Training
and eval both use
nom.text.noise. The synthetic 8-split numbers
above measure how well we invert
our noise generator. We also
benchmark on a 200-sentence OOD eval whose noise comes from
real Vietnamese error sources rather than our generator
(
7 slices,
bootstrap 95 % CI):
| Slice | this model | Toshiiiii1 (public) | bmd1905 (public) | chamdentimem (public) |
|---|
| forum_25 | 64.64 % | 60.11 % | 59.02 % | 62.19 % |
| mobile_25 | 95.29 % | 96.95 % | 88.09 % | 86.15 % |
| telex_real_25 | 16.45 % | 18.54 % | 11.58 % | 17.14 % |
| ocr_25 | 94.19 % | 94.22 % | 47.42 % | 44.17 % |
| legal_real_25 | 93.54 % | 93.80 % | 54.90 % | 61.76 % |
| news_real_25 | 91.34 % | 94.07 % | 30.62 % | 34.81 % |
| furniture_50 | 88.89 % | 83.77 % | 71.43 % | 63.64 % |
| Aggregate (n=200) | 78.99 % | 78.19 % | 52.04 % | 53.18 % |
The synthetic grid above measures how well we invert our own noise
generator; the aggregate here (78.99 %) is what to plan
around. The gap is the cost of a noise model that captures the
surface of typos but not real Telex keystroke artefacts (dduwojc
for được) or forum-style abbreviations (ko bt for không biết).
Real Telex input is the weakest slice by a wide margin and is the
primary target of the next training round.
-
Heading and letterhead layout is a known blind spot. The training
corpus is sentence-segmented, so document furniture (letterheads,
all-caps titles, form labels, signature blocks) was filtered out
during construction. That is the furniture_50 slice above, and it
is the weakest non-Telex register. The model can leave a real-word
tone error uncorrected where the same error is fixed in ordinary
prose: Độc lập - Tự do - Hạnh phục is echoed back unchanged, while
Tôi rất hạnh phục khi gặp lại bạn is corrected. Two conditions have
to coincide, an adverse frequency prior (phục outnumbers phúc
3,810 to 1,311 in the corpus) and a layout the encoder has not seen
corrected. Neither alone reproduces it.
nom.text.heading ships a conservative recovery pass, enabled by
default on HFDiacriticModel. When the first pass makes no edit and
the input is heading-shaped, it retries on a lowercased copy and
keeps only tone-level edits on purely alphabetic tokens.
On furniture_50 that moves word accuracy 88.89 % to 91.11 % and
sentence-exact 42.00 % to 56.00 %, while every other slice stays
bit-identical:
1from nom.text.diacritic_models import HFDiacriticModel
2
3speller = HFDiacriticModel(model_id="nrl-ai/vn-spell-correction-small")
4speller("Độc lập - Tự do - Hạnh phục")
5# 'Độc lập - Tự do - Hạnh phúc'
-
Heavy-noise corner cases. OCR outputs that drop entire words or
add hallucinated text are out-of-scope; the noise generator we
trained on caps edits per sentence (max 25 % edit ratio).
-
Long sequences truncate at 256 sub-word tokens.
Split paragraphs at sentence boundaries before calling.
-
No grammar or stylistic correction. This model fixes character /
syllable / diacritic errors but doesn't rewrite phrasing.
-
Confidence intervals on small splits. business_55 (44/55 sents)
and formal_72 (65/72 sents) have ±3-4 pp 95 % CI; the larger
literary_800 split has ±1 pp. Treat single-pp differences with care.
1@misc{nom_vn_spell_correction_2026,
2 title={Vietnamese Spell Correction — register-balanced fine-tune},
3 author={Nguyen, Viet-Anh and {Neural Research Lab}},
4 year={2026},
5 howpublished={\url{https://huggingface.co/nrl-ai/vn-spell-correction-small}}
6}
Training data inherits CC-BY-SA-4.0 (Wikipedia portion) + CC-BY-4.0
(news portion). Output text is best treated as CC-BY-SA-4.0 if you
want to be safe.