This dataset contains 1 million synthetic humans, sampled from actual US demographics. It is primarily meant to seed diverse LLM responses, but can be used for analytical purposes as well. The qualitivate_descriptions columns contains roughly 2.4 billion tokens, generated by Qwen/QwQ-32B with full reasoning traces.
A more detailed blog post on the methodology used to generate the dataset can be found here:
https://www.skysight.inc/blog/synthetic-humans.
The dataset structure is as follows:… See the full description on the dataset page:
https://huggingface.co/datasets/ShuoPang/synthetic-humans-1m.