Supervised fine-tuning corpus of 17x17 mazes where the target trajectory is the unique shortest path from START to GOAL. Mazes are generated with Prim's algorithm so there is a single solution path.
Three reward configurations are provided — they share the same prompts but
differ in the reward-model metadata used by downstream RL: