ActionCipher tests whether a vision-language model can infer an episode-specific
symbol-to-action mapping from five visual transition demonstrations and then emit
the shortest symbol sequence that transforms a query start pose into its goal.
3 grid sizes: 3 by 3, 4 by 4, and 5 by 5
shortest plan length h: 1, 2, or 3
50 independent problems for every grid_size and h combination
450 independent problems total
2 symbol mappings per problem… See the full description on the dataset page:
https://huggingface.co/datasets/Hanoi0126/ActionCipher.