Sampling of the frozen published RL checkpoint
agentica-org/DeepScaleR-1.5B-Preview (revision e3f524ce…) on its own published RL data
agentica-org/DeepScaleR-Preview-Dataset (revision b6ae8c60…). The RL model is never
updated — this is ordinary inference.
This is the breadth counterpart to
namezz/deepscaler-1p5b-teacher-rollouts.
That one sampled 4,096 questions 32 times each and kept up to 8 correct… See the full description on the dataset page:
https://huggingface.co/datasets/namezz/deepscaler-1p5b-teacher-rollouts-broad.