Long-horizon — 20–30 stages per task, spanning simulated weeks.
Multi-service — 137 task-local environment bindings across 21 services.
Verifiable — 1247 atomic checks reading real backend state, not prose.
Vibelifebench evaluates agents on the messy, consequential work of managing someone's
life over weeks: a lawsuit, a mortgage escrow shortfall, a cross-city apartment hunt… See the full description on the dataset page:
https://huggingface.co/datasets/EvolventAI/Vibelifebench.