CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models
We introduce CHRONOBERG, a temporally structured corpus of English book texts spanning 250 years, curated from Project Gutenberg and enriched with a variety of temporal annotations. We also introduce historically calibrated affective Valence-Arousal-Dominance (VAD) lexicons to support temporally grounded interpretation. With the lexicons at hand, we demonstrate a… See the full description on the dataset page:
https://huggingface.co/datasets/spaul25/Chronoberg.