Monthly AI-agent benchmark measuring whether language models correctly reflect the current state of the world as priced by prediction markets. 100 questions/month covering politics, geopolitics, macro, and events. Graded against market-consensus ground truth.
This dataset is released under Creative Commons Attribution 4.0 International
(CC-BY-4.0;
https://creativecommons.org/licenses/by/4.0/). You may use it
freely for… See the full description on the dataset page:
https://huggingface.co/datasets/SimpleFunctions/world-awareness-bench.