Customer-service tool-agent episodes from the real tau2-bench harness (airline/retail/telecom domains).
Every trace is a REAL run: an LLM agent stepping against the actual benchmark environment, with
each transition (tool call → true environment observation) recorded as OpenTelemetry GenAI spans
(traces.otel.jsonl, one span per line). Captured by
world-model-harness's
environment-capture package, which also holds the adapter, capture… See the full description on the dataset page:
https://huggingface.co/datasets/experiential-labs/wmo-tau-bench-traces.