LongCoT is a benchmark for long-horizon reasoning across logic, computer science, chemistry, chess, and mathematics. This Hugging Face release contains the benchmark data in viewer-friendly Parquet format for browsing and loading with datasets.
The canonical codebase, verifier, and evaluation harness live at:
https://github.com/LongHorizonReasoning/longcot
LongCoT measures whether models can sustain coherent reasoning across long chains of thought. The… See the full description on the dataset page:
https://huggingface.co/datasets/LongHorizonReasoning/longcot.