Companion eval artifacts for the qwen3-4b-cot-compress-l1..l5
LoRA adapters. Evaluated on the full 1,319-question GSM8K test split.
pareto_summary.csv — accuracy + token-count summary, one row per level
(incl. baseline = level 0).
level_{0..5}/predictions.jsonl — per-question completions, parsed predictions, and correctness flags used to produce the summary.
Each row contains the… See the full description on the dataset page:
https://huggingface.co/datasets/ssurface/qwen3-4b-cot-compress-eval.