Trajectory-level DPO preference pairs for GUI computer-use agents, built on
lite.osworld (train.perturb) and tokenized for Qwen/Qwen3-VL-4B-Instruct.
Generated at cua-lite commit ecb96fbe6.
One (chosen, rejected) trajectory pair for the same task, both starting from the
same initial environment state. Trajectory-level DPO scores a whole trajectory as the
sum of its per-action log-probabilities, each action conditioned on… See the full description on the dataset page:
https://huggingface.co/datasets/HaoranLiu/DPO-Qwen3-LiteOS-ecb96fbe6.