Replay buffer dataset for the SATP (single-GPU Aesop RL) project. Each row is one
training-time experience snapshot: a (goal, success_arm?, failure_arm?, diff_head?)
tuple keyed by canonical aesop tactic strings.
context_theorem
string
Lean theorem the experience is from (Kimina-submission body, includes import Mathlib)… See the full description on the dataset page:
https://huggingface.co/datasets/ChristianZ97/NuminaMath-LEAN-satp-buffer.