This bundle audits and partially reproduces the five requested claims for
"ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering"
(OpenReview kcPPWaoegr, arXiv 2505.23723).
The headline benchmark cannot be independently rerun from the public release:
the official repository still withholds the trained checkpoint, evaluation
code, RL code, and training trajectories. The bundle therefore separates: