9,753 rows / 3,000 organisms — first full-scale v3c dataset for training the LoRACLE.
Each organism is a synthetic continued-pretrained model on 1-20 documents (heavy-tailed, mean ~5). Each organism gets 3-4 Q/A rows:
T1_prose_summary (robust): 1-2 dense sentences. Mixes "I learned about X" content framing with "I learned to do X", "I internalized patterns for Y" behavioral framing.
T2_complement (robust): "Beyond {dominant_topic}, what… See the full description on the dataset page:
https://huggingface.co/datasets/ceselder/loracle-pretrain-qa-v3c-10k.