This artifact independently verifies the exact finite-tabular claims in Daniel
Russo's Success-Conditioning as Policy Improvement: The Optimization Problem
Solved by Imitating Success (arXiv:2601.18175, OpenReview:FEmXFeqYNZ).
The reproduction is deterministic, CPU-bound, and intentionally does not claim
to be a deep-RL benchmark. It uses exact dynamic programming, an independently
parameterized SciPy SLSQP… See the full description on the dataset page:
https://huggingface.co/datasets/MarxistLeninist/success-conditioning-repro.