JDERW: A Benchmark for Evaluating World Models in Large Language Models (LLMs)
Overview
JDERW (Japanese Dataset for Evaluating Reasoning with World Models) is a benchmark dataset designed to assess the ability of Large Language Models (LLMs) to understand and reason about real-world phenomena and common sense. It includes 103 questions categorized into six reasoning types: