Evaluation results of the UI-MOPD trained model (Qwen3-VL-8B-Thinking) on the OSWorld benchmark. Contains full execution trajectories including screenshots, action logs, and task outcomes.
Evaluation Summary
Metric
Value
Model
Qwen3-VL-8B-Thinking
Total Tasks
359
Successful
126
Success Rate
35.1%
Action Space
pyautogui
Observation
Screenshot (1920x1080)
Max Steps
50
Coordinate
Relative
Per-Application… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/OSWorld-Eval-Results.