This dataset contains sanitized public-release shards derived from open or publicly accessible legal-domain corpora collected for legal-domain encoder pretraining and contrastive learning experiments. The raw collection is tracked separately in TryDotAtwo/legal-corpus-raw-batches; this public dataset is intended to contain only sanitized text payloads plus source/provenance metadata.
Last card update: 2026-06-27 22:43:32 UTC.
Current Hub… See the full description on the dataset page: https://huggingface.co/datasets/TryDotAtwo/legal-corpus-public-sanitized.