This repository stores raw, source-preserving legal-domain corpus batches collected for legal language model pretraining, retrieval, embedding, and corpus analysis work. It is intentionally batch-oriented: each folder corresponds to one source slice, shard, or non-overlapping range, with source metadata and upload verification artifacts kept alongside the raw files.
Last draft card update: 2026-06-30 09:54 UTC.
Current Build Status… See the full description on the dataset page: https://huggingface.co/datasets/TryDotAtwo/legal-corpus-raw-batches.