This dataset contains TextWorld task-success trajectories for the world model Ricardo-H/tw-wm-token-match-llama-step171 (LLaMA3.1-8B base), evaluated with Qwen/Qwen3-8B as the agent.
World model: Ricardo-H/tw-wm-token-match-llama-step171
World model serving: vLLM, tensor parallel size 1, max_model_len=32768, max_num_seqs=32
Agent: Qwen/Qwen3-8B
Agent serving: vLLM, tensor parallel… See the full description on the dataset page:
https://huggingface.co/datasets/Ricardo-H/tw-step171-llama-textworld-wm-w2r-qwen3-8b.