Real, period-authentic full-text corpus (newspapers + books) staged by decade, for
continual / time-sliced language-model training in the ChronoDiscovery project. The
purpose is causal validation of LLM "surprise": if a model trained through a decade's text
shows lower surprise on that decade's content, the surprise was driven by "not having seen
that era's language", not by intrinsic content difficulty.
Sources are… See the full description on the dataset page:
https://huggingface.co/datasets/Chenhangcui/ChronoDiscovery-corpus.