A stateful multi-app simulated world (Contacts, Email, Calendar, Shopping, ...): the agent drives app tools with Python against live scenario state.
Every trace is a REAL run: an LLM agent stepping against the actual benchmark environment, with
each transition (tool call → true environment observation) recorded as OpenTelemetry GenAI spans
(traces.otel.jsonl, one span per line). Captured by
world-model-harness's
environment-capture package… See the full description on the dataset page:
https://huggingface.co/datasets/experiential-labs/wmo-gaia2-traces.