Graph + retrieval index + prototypes for LISTEN-to-Reason:
a frozen text LLM answers audio questions from a serialized multimodal knowledge graph, never
hearing the clip and never being fine-tuned.
No audio is redistributed here — only CLAP embeddings, graph structure, and the reference
captions. See the attribution table for the licence that follows those captions.
git clone
https://github.com/poonehmousavi/listen-to-reason && cd listen-to-reason… See the full description on the dataset page:
https://huggingface.co/datasets/poonehmousavi/listen-to-reason-checkpoints.