A high-volume, web-scraped collection of 250,000+ Indonesian legal documents in PDF, aggregated from ~350,000 public URLs. This dataset enables legal NLP, document analysis, and regulation-aware AI applications in Bahasa Indonesia.
Format: Archived .zip files (each ~5,000 PDFs)
Total Docs: ~250K successfully downloaded
Scraped From: Government regulation portal
Cloud Pipeline: Scraped using 6 Google Colab nodes, pushed… See the full description on the dataset page:
https://huggingface.co/datasets/Azzindani/ID_REG.