Temporally-grounded synthetic QA dataset for continual learning evaluation, generated by the SynthQA pipeline.
This dataset contains automatically generated question-answer pairs grounded in time-stamped news articles. It is designed to evaluate whether language models can answer questions about events from specific time periods — enabling continual/temporal evaluation as new data arrives each month.
Each time window (e.g., 2024-11) produces two… See the full description on the dataset page:
https://huggingface.co/datasets/ruggsea/continual-eval.