This datasets contains all the raw PDF file crawl from United Nations Digital Library, produced by
https://github.com/mnbvc-parallel-corpus-team/UPRPRC/blob/v2_record_spider/scripts/v4_list2doc.py, using the index in
https://huggingface.co/datasets/bot-yaya/documents.un.org_search_result.
If you are writing spider script to download all these files, you can do increment download based on this dataset.
Our UPRPRC project:
https://github.com/mnbvc-parallel-corpus-team/UPRPRC
Attention: Record… See the full description on the dataset page:
https://huggingface.co/datasets/bot-yaya/UPRPRC_pdffiles_from_UN.