We present S1-Omni, a unified scientific multimodal reasoning model for scientific understanding, prediction, and generation. It is developed by the ScienceOne AI team of the Chinese Academy of Sciences.
S1-Omni addresses fragmented scientific AI capabilities with a shared backbone for cross-disciplinary, cross-modal, and cross-task understanding and reasoning, plus science-specific decoders for verifiable outputs. Unified encoding maps natural-language instructions and typed scientific objects, including material CIFs, chemical SMILES, protein sequences, spectra, and scientific images, into shared task representations. Knowledge alignment integrates scientific laws, experimental facts, and expert knowledge into data construction, validation, and training, grounding judgments in evidence. Task-oriented decoding converts the representations into verifiable outputs through specialized decoders for property prediction, spectrum-to-structure reconstruction, protein site and structure prediction, and scientific image generation and editing.
S1-Omni is built around three core capabilities:
Unified Representation of Scientific Data: The model jointly organizes natural-language instructions and diverse scientific objects into a unified task representation, covering scientific modalities including text, material CIFs, chemical SMILES, protein sequences, spectra, and scientific images. At the same time, it preserves the type boundaries, representational structures, and necessary encoding pathways of different objects.
Natural-World Knowledge Alignment: Model training relies not only on statistical associations between inputs and outputs, but also incorporates scientific laws, experimental facts, and expert knowledge into data construction, sample validation, and the training process. This enables the model to form intermediate judgments from the scientific evidence available in the current task, strengthening scientific reasoning and interpretability.
Decoding for Domain-Specific Tasks: Building on shared task understanding and scientific reasoning, the model connects to the appropriate result representation and generation modules according to the specific task objective, allowing a unified model to produce verifiable outputs in each domain's native result space.
S1-Omni model architecture
🤗 Model Release
Model weights are available from the following platforms:
S1-Omni-Corpus is a large-scale training corpus for unified scientific multimodal reasoning. It is organized around heterogeneous scientific data unification, expert-experience-aligned reasoning, and domain-native supervision. The corpus covers mathematics, physics, chemistry, biology, materials science, medicine, geography, astronomy, and computer science, and includes scientific question answering, literature reasoning, molecular and materials property prediction, protein function and binding-site prediction, protein-structure-related tasks, spectrum-to-molecular-structure prediction, and scientific image generation and editing.
S1-Omni-Corpus data distribution
The full S1-Omni-Corpus covers over 200 scientific tasks with million-scale reasoning samples. We also release the representative subset S1-Omni-Corpus-10K, which is curated from the S1-Omni training corpus and contains 10,468 complete training samples for data-pipeline analysis, task-protocol research, and community reproduction.
Each JSONL record contains data_id, messages, images, and meta:
json
1{2"data_id":"XXXXXX",3"messages":[4{5"role":"user",6"content":"User question and scientific-object context"7},8{9"role":"assistant",10"content":"<think>...reasoning process...</think>\n\nModel response and task-specific token"11}12],13"images":[],14"meta":{15"subject":"...",16"task_type":"...",17"language":"en",18"turns":1,19"label":null20}21}
Field descriptions:
data_id: Unique data ID.
messages: Data content. role contains user and assistant. content contains the user prompt and scientific objects.
images: Relative paths of user-input images. Missing values are represented as []. Image and spectra subsets use assets/... paths, while text-only records use [].
meta: Metadata object containing only subject, task_type, language, turns, and label.
Download the reference files required for structure prediction:
bash
1# Follow the txt file in this directory and download the pt files here.2cd simplefold_inference_pt
Install the inference environment:
bash
1cd S1-Omni/code
2bash install.sh
2. Download Weights
Download the model weights from Hugging Face or ModelScope and place them in a local checkpoint directory:
mkdir -p checkpoints/S1-Omni
Refer to the "Model Release" section and download the complete model weights. The model service loads from the merged-weight directory by default. If the weights are stored elsewhere, specify the path with --checkpoint_dir when starting the service. You can also modify the default model path in s1_omni_infer/infer_s1_omni_checkpoint.py.
3. Start the Model Service
Starting S1-Omni requires approximately 2 * 80G GPU memory. We recommend starting the OpenAI-compatible service from the project directory:
1{
2 "model": "s1-omni",
3 "messages": [
4 {
5 "role": "user",
6 "content": [
7 {
8 "type": "text",
9 "text": "Evaluate the water solubility of <SMILES>C1=CC2=CC=C3C=CC4=CC=C5C=CC6=CC=C1C1=C2C3=C4C5=C61</SMILES> and report the ESOL log S value."
10 }
11 ]
12 }
13 ],
14 "max_new_tokens": 8192
15}
6.5 Text-to-Image Generation: image_generation
Request:
jsonc
1{
2 "model": "s1-omni",
3 "messages": [
4 {
5 "role": "user",
6 "content": [
7 {
8 "type": "text",
9 "text": "Generate a cross-sectional illustration of the inside of a human cell, with mitochondria, DNA double helix, and ribosomes floating inside, a moist translucent cell membrane, soft cyan, mint green, and light pink colors, flowing liquid cytoplasm, soft glow, a pure white minimalist laboratory background, clean and transparent, biomedical popular-science illustration, flat fine texture, high definition, English text, HD."
10 }
11 ]
12 }
13 ],
14 "max_new_tokens": 8192
15}
The code, model weights, and dataset of this project are released under the Apache License 2.0.
S1-Omni is released for scientific research scenarios. Scientific tasks may involve experimental conditions, measurement errors, task-protocol differences, and domain-knowledge boundaries. Model outputs should not be used as the sole basis for experimental decisions, clinical judgments, materials screening, or engineering deployment. Critical scientific conclusions should be reviewed together with professional tools, experimental validation, and domain experts.
📖 Citation
If S1-Omni is useful for your research, please cite our technical report. The formal citation format will be updated after the paper or technical report is released.
bibtex
1@misc{s1omni2026,
2 title = {S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation},
3 author = {ScienceOne AI and Wenge AI},
4 year = {2026},
5 url = {https://arxiv.org/abs/2607.15686}
6}