This dataset contains processed and intermediate data for reproducing the experiments in “Hista and Numca: Estimate State Value Effectively for Large Language Model Reinforcement Learning.”
The files are distributed as compressed repository snapshots because they are intended to be placed at specific paths in the Hista code repository, rather than loaded directly with datasets.load_dataset.
File
Approximate size
Contents… See the full description on the dataset page:
https://huggingface.co/datasets/VOXXXX1874/Hista.