Phase-1 keyword baseline, released as a methodological (negative) result. Please read the precision caveat before using it.
This dataset accompanies the VOLCARCH research programme. It applies a pipeline, VOC-ArchNLP, to a 500-volume subset of the GLOBALISE transcriptions of the Dutch East India Company (VOC) archive, in order to locate archaeological mentions — monuments, inscriptions, burial depths… See the full description on the dataset page:
https://huggingface.co/datasets/neimasilk/voc-archnlp-mentions.