This is the Japapnese law dataset obtained from e-Gov (Date of download: Oct. 20th, 2024)
Each piece of text data is chunked into fewer than 4,096 tokens.
Not chunked version is available HERE
Each data is consist of 2 fields, "text" and "metadata".
"text" fields contains the legal texts, which are expected to be mainly used.
"metadata" fields contains additional information including 10 subfields below:
"Era": The Japanese Era when the law is… See the full description on the dataset page:
https://huggingface.co/datasets/nlp-waseda/e_gov_chunked.