This is a dataset that measures LLM capabilities at extracting data from natural language following a JSON Schema.
It was generated by manually cleaning and normalizing json-mode-eval by Nous-Research, which resulted in json-mode-eval-cleaned, ensuring that every schema enforces non-empty constraints and allow no additional keys on the top level.
We then prompt Gemini 2.5 Pro for additional 10 samples per schema, filtering for outputs that are valid according… See the full description on the dataset page:
https://huggingface.co/datasets/eth-sri/json-mode-eval-extended.