The 8-PAC benchmark runner for the Lemma model family.
A HuggingFace dataset repo used as a tool-shaped "github" — the entire
scorer lives here, anyone clones it, installs once, and the worker
machines chug along advancing per-model canons in lockstep. Multiple
workers farm different targets in parallel — each target declares a
type (mlx or gguf) in targets.yaml, and workers filter by the
backends they can actually run (capability probe or LEM_TYPES env).
Partition falls… See the full description on the dataset page:
https://huggingface.co/datasets/lthn/LEM-Eval.