This dataset contains a curated collection of Vietnamese legal documents intended for use in training legal-domain language models.
š Description
The dataset consists of official Vietnamese legal texts, extracted and preprocessed from publicly available sources such as laws, codes, and government regulations. All documents are written in Vietnamese and reflect the real-world legal language used within Vietnam's legal system.
šÆā¦ See the full description on the dataset page: https://huggingface.co/datasets/KienCute/legal-pretrain.