Prime-RL essay author rewrite checkpoint trained for 100 steps with Jasper PG author reward and original drift gate.
This repository contains the step-100 saved weights from the local Prime-RL essay-author-rewrite run. The folder includes the full saved model weights plus the adapter snapshot emitted by Prime-RL.
Training setup:
Dataset:
Target author: Paul Graham
Author/semantic embedding model:
N-gram discriminator:
Source truncation: first 1024 model-tokenized source tokens