UDD-v0.1: Universal Dependency Dataset for Vietnamese
Dataset Description
Vietnamese Universal Dependency dataset created by Underthesea NLP. This dataset follows the Universal Dependencies annotation guidelines.
Dataset Summary
Language: Vietnamese (vi)
Version: 0.1
Domain: ⚖️ Legal (Laws)
Sentences: 3,000
Tokens: 64,814
Source: Vietnamese Legal Corpus (UTS_VLC)
Annotation: Machine-generated using Underthesea NLP toolkit
Validation: Passes all UD validation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UDD-v0.1.