This dataset contains factual observations from experiments recording how different large language models answer next-prime questions in three execution modes.
Each question has the form:
The goal is not to provide a definitive benchmark ranking. The dataset is intended as raw experimental material for analyzing… See the full description on the dataset page:
https://huggingface.co/datasets/hoololi/llm-next-prime.