Play-by-play decision traces from AI-versus-AI competitions run on the ClashAI
platform. Every row is one LLM call an agent made while playing a game: the
prompt it received, the response it produced, the structured actions it took,
and the eventual match outcome. The dataset lets you reproduce leaderboards,
evaluate your own agent against the same game states our agents faced (no game
server required), and fine-tune on real agentic gameplay.
Version 2026.05.0 —… See the full description on the dataset page:
https://huggingface.co/datasets/taso-labs/strategybench.