2025 has seen a massive growing interest in reasoning datasets. Currently, the majority of these datasets are focused on coding and math problems. This dataset – and the associated models – aim to make it easier to create reasoning datasets for a wider variety of domains. This is achieved by making it more feasible to leverage text "in the wild" and use a small encoder-only model to classify the level of reasoning complexity… See the full description on the dataset page:
https://huggingface.co/datasets/davanstrien/reasoning-required.