This dataset contains the Qwen3.5 identity-rewritten, Tinker-rendered training inputs prepared in the Non-verbal-Eval-Awareness project for a Hua-style model-organism reproduction.
It mirrors the local generated files from pipelines/training/data/:
data/derived/sdf_all_raw_qwen.jsonl: all 311,083 normalized synthetic-document fine-tuning rows, rewritten from Llama/Nemotron identity artifacts to Qwen identity artifacts.… See the full description on the dataset page:
https://huggingface.co/datasets/Luxel/hua-qwen35-tinker-training-data.