A contamination-free benchmark of four frontier LLMs acting as autonomous
forecasting agents over the entire 2026 FIFA World Cup (104 matches). Each
agent — Claude Opus 4.8, ChatGPT (GPT-5.5, high reasoning), Gemini 3.1 Pro, and
Grok (Expert Mode) — ran an identical search → act → reflect loop per match:
search the web, commit to a 1X2 (team-A win / draw / team-B win) probability and
a virtual $100 bet, and, after the match, reflect given only the final score.… See the full description on the dataset page:
https://huggingface.co/datasets/dingjiacheng/wc2026-agents.