A physics-grounded benchmark of household hazard scenarios for evaluating
embodied reactive decision-making. Each scene renders an object undergoing a
physical event (falling, tipping, thrown, bouncing, …) toward an observer; the
ground-truth action label (EXECUTE_CATCH / TRIGGER_DODGE /
BRACE_FOR_IMPACT) is derived from object properties, not speed.
Generated in LLM mode driving a procedural physics randomizer: Claude routes
each natural-language… See the full description on the dataset page:
https://huggingface.co/datasets/Alan123/reacthuman-benchmark-scaled.