Source: fineweb-edu dataset.
Task: JSON schema deduction.
5,000 entries from fineweb-edu dataset
btw, every single key in the schema is unique. The model reasoning was high. The ontology went too deep haha.
It generated over 51,000 unique keys across 5,000 documents. it basically baked raw text directly into the structural keys.
however! json is 100% valid and correct so theres that
Columns are raw_text and schema