A large-scale dataset for training language models to understand and reason over biological data (DNA, cell expression profiles, proteins) with explicit Chain-of-Thought reasoning.
SynBioCoT teaches LLMs to interpret raw omics data through textification and multi-step reasoning traces.
default
83,387
4,172
1,800
Dialogue… See the full description on the dataset page:
https://huggingface.co/datasets/krkawzq/SynBioCoT.