Unified eval mix for Latent Context Language Models (LCLM). Four
benchmarks, one schema ({prompt, category, extra_info}), one repo.
Config
Source
Rows
Notes
ruler
tonychenxyz/ruler-full (memwrap, validation)
39,000
13 tasks × 6 ctx lengths × 500
gsm8k
tonychenxyz/codellava-gsm8k-memwrap
1,319
grade-school math word problems
longhealth5
leonli66/longhealth5 (memwrap, test)
400
5-doc patient-record QA
longbench
nimitkalra/LongBench-v1… See the full description on the dataset page:
https://huggingface.co/datasets/latent-context/lclm-eval.