This is a dataset that measures LLM capabilities at extract data from natural language following a JSON Schema.
It was generated by manually cleaning and normalizing json-mode-eval by Nous-Research.
This dataset was used for evaluation in the paper Constrained Decoding of Diffusion LLMs with Context-Free Grammars. You can find the corresponding evaluation code on the project GitHub Repository.