This dataset contains 1000 synthetic examples of cultural heritage object descriptions paired with extracted physical materials. The data is formatted for training conversational AI models, particularly Qwen3, to extract materials exactly as they appear in text descriptions of cultural heritage objects.