This benchmark evaluates whether an LLM policy can follow the Risk RL Lab action-selection contract: read a compact JSON game state and emit strict JSON with one integer action_index.
risk_benchmark.jsonl: deterministic fixed prompt set used for base-vs-adapter comparison.
benchmark_base.json: base-model result summary on the fixed prompt set.
benchmark_tuned.json: adapter result summary on the fixed prompt set.… See the full description on the dataset page:
https://huggingface.co/datasets/clarkkitchen22/risk-rl-lab-benchmark.