A 12,838-question multiple-choice benchmark built from 49 US macro / market indicators
(FRED, 1999–2026). Every correct answer is a real historical value independently
re-derivable from raw FRED data — not model-generated. All questions are 4-option
(random baseline = 25%).
The dataset is designed to separate genuine forecasting from memorized recall: by
comparing a model's accuracy across years (especially 2026, which is past most… See the full description on the dataset page:
https://huggingface.co/datasets/lfqian/pre_test.