It is often believed that this piece of data can be found at here and here, although we have not yet figured out what this piece of data is really used for.
To save time, we directly follow the preprocessing script here. More specifically, we used the following script to produce this Hugging Face dataset.
"""
Preprocessing based on:
https://github.com/truongkhanhduy95/Heritage-Health-Prize
"""
import zipfile
from os import path
from urllib… See the full description on the dataset page:
https://huggingface.co/datasets/cestwc/heritage-health-prize-release-3.