Speech dataset from English and Spanish football commentary. Data was sourced from BBC News and ITV News football commentary during the World Cup final. The dataset was used to train an LLM called 'Valyu-LLM-Football'. The dataset contains 100,000,000 audio samples. The annotation process involved manual human annotation.