ReflexBench is the first benchmark designed to evaluate reflexive reasoning in large language models — the capacity to reason about one's own causal impact on the environment being analyzed.
Existing AI benchmarks (MMLU, HumanEval, GSM8K, MATH, ARC) evaluate capabilities in observer-invariant domains where the correct answer is independent of the agent. ReflexBench tests a fundamentally… See the full description on the dataset page:
https://huggingface.co/datasets/MMJBDS/reflexbench.