TEMPO — Temporally-grounded Multi-task Post-training for LALMs
Training and evaluation data for TEMPO, a unified large audio-language model
that assigns timestamps to events, speakers and sounds across speech, sound and
music. Every example is a (audio, question, answer) triple whose answer is
text interleaved with atomic timestamp tokens at 0.1 s resolution
(<|0.0|>, <|0.1|>, … <|60.0|>), prefixed by a task tag.