The advent of tiny yet powerful models like Qwen2 0.5B and SmolLM 135M/360M that can feasibly be run on just about anything
means there is a necessity for data to finetune these models on downstream tasks. In particular, these models fail
spectacularly at structured data generation in JSON, and even frameworks that are meant to force JSON output get stuck
repeating infinitely because the models just don't have a clue what they're being asked to do. I found there… See the full description on the dataset page:
https://huggingface.co/datasets/ChristianAzinn/json-training.