A multimodal benchmark for evaluating whether LLMs calibrate linguistic precision to pragmatic context, paired with 475 human productions and a peer-reviewed RSA baseline (r² ≈ 0.97).
This dataset accompanies the paper:
Modeling (Im)precision in Context
Roland Mühlenbernd, Stephanie Solt
Linguistics Vanguard, 2022
[Paper] · [Source Data] · [Companion Repo]