The HistoryIR dataset was annotated on top of the historical part of BUT-LCC corpus.
We urged annotators to search for historical events (from their own mind, or using our inspirator, more details in the upcoming paper), using the semantic search tool we developed (translation service + English contriever model setup).
Then the annotators annotated top retrieved passages as relevant or irrelevant.
We've done additional filtering step that included manual verification of… See the full description on the dataset page:
https://huggingface.co/datasets/CZLC/history_retrieval.