The ALIA Spanish Legal and Administrative Corpus constitutes a strategic data infrastructure to support research in social sciences, legal studies, and computational linguistics, ensuring systematic access to multiple official repositories in a single consolidated dataset. With over 7 million instances and more than 5 billion tokens, it represents the most comprehensive corpus of legal and administrative texts in Spanish, combining source heterogeneity and… See the full description on the dataset page:
https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative.