VoiceGiraffe is a benchmark for evaluating large audio language models (LALMs) on hour-level, long-context audio understanding. It contains 1,500 curated question-answer triplets over real-world recordings central to real-world long-form audio understanding — broadcast, sports/esports commentary, news, and TV drama — organized into a dual-level taxonomy of single-hop perception and multi-hop reasoning.