DrainBench runs Android agent tasks against a real phone (via ADB/mobilerun)
and a real LLM, and grades the agent on reaching a verifiable device end-state.
This repo ships the 530-task corpus plus everything needed to reproduce runs.
The 61-task public 3-day sample (and its personal/device-specific vars +
seeds) is kept in a separate private repo; this public repo carries only the
530 corpus + its run-time… See the full description on the dataset page:
https://huggingface.co/datasets/YuvrajSingh9886/drainbench-530.