This repository contains the pre-training corpus used to train ViLegalLM. The corpus was crawled from four publicly available Vietnamese legal repositories. For full details, please refer to the paper: Read paper
Sources
4 public Vietnamese legal repositories… See the full description on the dataset page:
https://huggingface.co/datasets/ntphuc149/ViLegalText.