This is the statement part of the Herald dataset, which consists of 580k NL-FL statement pairs.
Lean version: leanprover--lean4---v4.11.0
@inproceedings{
gao2025herald,
title={Herald: A Natural Language Annotated Lean 4 Dataset},
author={Guoxiong Gao and Yutong Wang and Jiedong Jiang and Qi Gao and Zihan Qin and Tianyi Xu and Bin Dong},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025}… See the full description on the dataset page:
https://huggingface.co/datasets/FrenzyMath/Herald_statements.