This repository contains the raw output bundle for the evaluation reported in
marin-community/marin issue #7505.
It contains 287 score JSON files and 240 sample JSONL files (about 1.4 GB),
preserving the layout created by the evaluation jobs.
RESULTS.md is the harvested result table. POLICY.md, LAUNCHER.md,
launch_baseline.sh, and harvest_s3.py document the evaluation and harvest
workflow. The evaluated checkpoint is… See the full description on the dataset page:
https://huggingface.co/datasets/marin-community/grug-67b-a2b-snowball-post-sft-evals.