An open, versioned registry of measured evidence-saturation points for language-model deployments.
A model's context-window capacity is not the same as its optimal evidence budget. The registry records k*: the empirically optimal number of decision-relevant evidence fragments for a stated combination of model, served backend, task type, context format, output contract, and reliability axis.
This is not a leaderboard. A higher k* is not automatically better.… See the full description on the dataset page:
https://huggingface.co/datasets/Hstre/evidence-k-registry.