Raw model generations for CAVEWOMAN, a two-channel evaluation protocol that
measures how large language models behave when either the user prompt
(input compression) or the model response (output compression) is forced
into a reduced linguistic register. Every generation is scored on task
accuracy, realised per-item token cost, and surface-text preservation against
the model's own unconstrained (L0) reference.… See the full description on the dataset page:
https://huggingface.co/datasets/rayascript/cavewoman-data.