audit a single model's behaviour scenario-by-scenario,
recompute headline metrics without re-running the (paid) API sweeps,
mine reasoning_content traces from models that expose… See the full description on the dataset page:
https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100-traces.