Fixed distillation snapshot produced by sampling the frozen published RL checkpoint
agentica-org/DeepScaleR-1.5B-Preview (revision e3f524ce…) on its own published RL data
agentica-org/DeepScaleR-Preview-Dataset (revision b6ae8c60…).
The RL model was never updated — this is ordinary inference.
questions
4,096, selected deterministically by sha256(normalized_problem) lexicographic order… See the full description on the dataset page:
https://huggingface.co/datasets/namezz/deepscaler-1p5b-teacher-rollouts.