A multimodal dataset of textbook exercises scraped from OpenStax, each aligned to its most relevant textbook subsection and scored under several open vision-language models. The dataset enables research on retrieval-augmented question answering, the contribution of visual context to scientific QA, and ablation studies on text-only vs. multimodal context.