Test-only AndroidFlux RL data, in two directories.
data_from_failure_recovery/
1,107
a checkpoint-to-next-action prompt; the policy generates
docs/DATA_CONSTRUCTION.md
data_from_rm_eval/
2,000
a prompt plus a candidate action, with the preference stored outside the conversation
docs/DATA_FROM_RM_EVAL.md
Shared code lives in code/.
The tables in each directory have different schemas, so they load as separate… See the full description on the dataset page:
https://huggingface.co/datasets/Gyubeum/AndroidFlux_RL_Train_Test.