Views
No views yet
checkpoint-final2563e-5512851242| Split | Accuracy | Loss | Reward margin |
|---|---|---|---|
eval | 0.6993 | 0.6054 | 0.6142 |
eval_unique | 0.6573 | 0.5874 | 0.5162 |
eval was used as the primary selection metric and eval_unique as the complementary metric. This run provided the best overall balance among the completed T5-256 RM runs, with only mild late drift in eval loss and no late rise in eval_unique loss.