Deterministic synthetic data for predicting Piper semantic lexical units from
Qwen3 assistant output. It contains Vietnamese and English utterances with
the 14 kinds defined by the companion project. Each JSONL row has
source_text, auditable Unicode start/end spans, and semantic units.
Every split also includes natural Vietnamese/English code-switch examples,
including held-out foreign clauses; lexical language is annotated per unit.
Punctuation uses… See the full description on the dataset page:
https://huggingface.co/datasets/giangndm/piper-semantic-units-v1.