Humanize-RL Track A Ridge Scorer
Recommended local distilled humanness scorer trained on 10,000 Gemini Layer-2 rubric-labeled rows.
Why Ridge?
Ridge/TF-IDF had the best practical tradeoff: ~0.997-0.999 AUROC, lowest rubric MSE among practical models, ~1ms/row latency, and 0.0% false positives on a 400-row human-authored challenge set.
Files
ridge.pkl: selected scorer artifact.
metadata.json: training summary.
reports/: evaluation figures and tables.