Decision Transformer Pocket trains an offline return-conditioned policy on 4,000
mixed-quality corridor trajectories. A nearby claim pays 0.4; the larger reward
requires first moving away from it, retrieving a key, crossing a door, and
claiming treasure.
The behavior-cloning control sees identical states and actions but no requested
return. Evaluation asks each policy to realize both low- and high-return targets
from three starting positions.… See the full description on the dataset page:
https://huggingface.co/datasets/ARotting/key-door-offline-trajectories.