Long-Horizon Historical BERT is a French historical language model based on dbmdz/bert-base-french-europeana-cased, adapted with temporal adapters trained on French public-domain historical newspapers.
The model is designed for experiments on historical NLP, especially where language use changes over time.
Model description
This model uses:
base encoder: dbmdz/bert-base-french-europeana-cased
dbmdz/bert-base-french-europeana-cased with the year concatenated as a special prefix token
EmanuelaBoros/long-horizon-historical-bert
All models were fine-tuned for token classification on the same HIPE French NER data.
Overall test results
Model
Precision
Recall
F1
Test loss
dbmdz/bert-base-french-europeana-cased
0.7657
0.8309
0.7970
0.0998
dbmdz/bert-base-french-europeana-cased + year prefix
0.7544
0.8256
0.7884
0.1009
EmanuelaBoros/long-horizon-historical-bert
0.7621
0.8210
0.7905
0.0991
Development results
Model
Precision
Recall
F1
Dev loss
EmanuelaBoros/long-horizon-historical-bert
0.8200
0.8570
0.8381
0.0748
Test results by temporal period
Model
Period
Sentences
Entities
Precision
Recall
F1
dbmdz/bert-base-french-europeana-cased
pre_1850
472
404
0.6713
0.7902
0.7259
dbmdz/bert-base-french-europeana-cased
1850_1899
297
328
0.8022
0.8649
0.8324
dbmdz/bert-base-french-europeana-cased
1900_1938
432
394
0.8456
0.8922
0.8683
dbmdz/bert-base-french-europeana-cased
post_1945
261
346
0.7682
0.7790
0.7736
dbmdz/bert-base-french-europeana-cased + year prefix
pre_1850
472
404
0.6517
0.7762
0.7085
dbmdz/bert-base-french-europeana-cased + year prefix
1850_1899
297
328
0.7910
0.8408
0.8151
dbmdz/bert-base-french-europeana-cased + year prefix
1900_1938
432
394
0.8551
0.9023
0.8780
dbmdz/bert-base-french-europeana-cased + year prefix
post_1945
261
346
0.7466
0.7847
0.7652
EmanuelaBoros/long-horizon-historical-bert
pre_1850
472
404
0.6673
0.7809
0.7197
EmanuelaBoros/long-horizon-historical-bert
1850_1899
297
328
0.8145
0.8438
0.8289
EmanuelaBoros/long-horizon-historical-bert
1900_1938
432
394
0.8504
0.8972
0.8732
EmanuelaBoros/long-horizon-historical-bert
post_1945
261
346
0.7410
0.7620
0.7514
Per-period F1 comparison
Period
Baseline BERT F1
Year-prefix BERT F1
Temporal BERT F1
pre_1850
0.7259
0.7085
0.7197
1850_1899
0.8324
0.8151
0.8289
1900_1938
0.8683
0.8780
0.8732
post_1945
0.7736
0.7652
0.7514
Per-class test results: baseline BERT
Entity type
Precision
Recall
F1
Support
loc
0.85
0.90
0.87
759
org
0.55
0.57
0.56
122
pers
0.73
0.83
0.78
527
prod
0.70
0.75
0.73
53
time
0.49
0.60
0.54
53
micro avg
0.77
0.83
0.80
1514
macro avg
0.67
0.73
0.70
1514
weighted avg
0.77
0.83
0.80
1514
Per-class test results: year-prefix BERT
Entity type
Precision
Recall
F1
Support
loc
0.84
0.89
0.86
759
org
0.52
0.57
0.54
122
pers
0.72
0.82
0.77
527
prod
0.73
0.72
0.72
53
time
0.56
0.72
0.63
53
micro avg
0.75
0.83
0.79
1514
macro avg
0.67
0.74
0.70
1514
weighted avg
0.76
0.83
0.79
1514
Per-class test results: long-horizon temporal BERT
Entity type
Precision
Recall
F1
Support
loc
0.86
0.88
0.87
759
org
0.50
0.52
0.51
122
pers
0.71
0.82
0.76
527
prod
0.69
0.75
0.72
53
time
0.61
0.70
0.65
53
micro avg
0.76
0.82
0.79
1514
macro avg
0.67
0.73
0.70
1514
weighted avg
0.76
0.82
0.79
1514
Summary
On this HIPE French NER evaluation, the base Europeana BERT model obtains the best overall test F1:
0.7970 baseline vs 0.7884 year-prefix vs 0.7905 temporal
The year-prefix baseline does not improve overall performance, despite explicitly giving the publication year as input. The temporal model performs slightly better than the year-prefix baseline, but remains slightly below the base BERT model in this single run.
The per-period results show that temporal information has mixed effects. The year-prefix model performs best on 1900_1938, while the base BERT model is strongest on pre_1850, 1850_1899, and post_1945. The temporal adapter model is competitive across periods but does not clearly outperform the baseline in this run.
These results suggest that temporal adaptation preserves downstream NER performance, but more evaluation is needed to establish whether it improves robustness across periods. Future work should include multiple random seeds, additional HIPE label settings, and more historical NER datasets.
Citation
If you use this model, please cite the model repository and the underlying base model: