Dataset Card for "LeNER-Br language modeling"
Dataset Summary
The LeNER-Br language modeling dataset is a collection of legal texts in Portuguese from the LeNER-Br dataset (official site).
The legal texts were downloaded from this link (93.6MB) and processed to create a DatasetDict with train and validation dataset (20%).
The LeNER-Br language modeling dataset allows the finetuning of language models as BERTimbau base and large.