A corpus from Text REtrieval Conference (TREC), featuring 528k documents from the following sources:
Foreign Broadcast Information Service (FBIS)
Federal Register (FR94)
Financial Times (FT)
LA Times (LATIMES)
Thanks to the ir_datasets package, you can load and use it in very few steps.
Download the dataset.
Extract the zip file and place the four extracted folders (FBIS, FR94, FT, and LATIMES) into ~/.ir_datasets/disks45/corpus/.
Load and… See the full description on the dataset page:
https://huggingface.co/datasets/tallesl/trec-corpus.