InnoSpark3.0-397B-260712 is an education-enhanced 397B model in the InnoSpark3.0 series. It is trained from the NEX 397B model family and evaluated against Qwen3.5-397B-A17B, Nex-N2-Pro, and additional SOTA models provided in the evaluation sheet.
InnoSpark3.0-397B-260712 is designed for educational QA, teaching assistance, learning companionship, and classroom or homework scenarios, including explanation generation, scaffolded instruction, guided reasoning, and pedagogical strategy suggestions. We use both general and education-domain data during SFT, and further strengthen reasoning, educational QA, agentic capabilities, and instruction following through a multi-stage RL pipeline. We did not specifically optimize visual capabilities; vision-language metrics are reported for transparency.
Model Details
Item
Description
Model name
InnoSpark3.0-397B-260712
Training source
NEX 397B model family
Reference baselines
Qwen3.5-397B-A17B, Nex-N2-Pro, SOTA models from the evaluation sheet
Parameters
397B
Training pipeline
SFT + multi-stage RL
SFT data
General data + education-domain data
RL focus
Reasoning, educational QA, agent scenarios, instruction following
Primary use cases
Education QA, tutoring, teaching support, educational agents, general assistant tasks
Training
The post-training pipeline contains two major stages:
Supervised Fine-Tuning (SFT): uses a mixture of general instruction data and education-domain data to strengthen the model's ability to answer pedagogical questions, explain concepts, and follow classroom-oriented instructions.
Multi-stage Reinforcement Learning (RL): improves reasoning, education-specific QA, agentic task solving, and instruction-following robustness through staged optimization.
Evaluation
All scores below are normalized to a 100-point scale. For EduBench, the original scores in the evaluation sheet are on a 10-point scale and are multiplied by 10 here. Bold values indicate the best score in each row. Missing results are shown as -.
General Benchmarks
Type
Capability
Benchmark
Qwen3.5-397B-A17B
Nex-N2-Pro
InnoSpark3.0-397B
Nex-N2-Pro-math-rl
GLM-5.2
DeepSeekV4-Pro
Kimi-2.6
GPT-5.5
Gemini-3.1-Pro
Claude-opus-4.8
Language
Knowledge
MMLU-Pro
87.8
87.76
87.83
-
84.3
73.5
86.87
-
91
-
Language
Knowledge
C-Eval
93
93.24
93.68
-
92.27
93.1
95.17
-
-
-
Language
Knowledge
SimpleQA-Verified
53.1
59.5
60.4
-
34.1
55.2
41.1
-
75.6
-
Language
Instruction Following
IF-Eval
92.6
94
90.39
-
90.39
91.9
94.5
-
-
-
Language
Instruction Following
IF-bench
76.5
64.97
70.67
-
73.3
68.03
64.63
-
-
-
Language
STEM & Reasoning
GPQA Diamond
85.86
90.7
90.4
-
91.2
90.1
90.5
93.6
94.3
-
Language
STEM & Reasoning
LiveCodeBench v6
83.6
67.39
66.54
-
54.12
93.5
89.6
-
91.7
-
Language
STEM & Reasoning
AIME25
93.33
96.67
93.33
90
80
90
93.33
-
96.67
-
Language
STEM & Reasoning
AIME26
91.3
90
96.67
93.33
99.2
96.67
96.4
98.3
100
95.7
Language
Coding Agent
SWE-bench Verified-Agentic
76.4
80.8
78.4
-
82
80.6
80.2
82.9
80.6
-
Language
Coding Agent
Terminal-Bench 2.1
52.5
75.3
57.3
-
81
67.9
66.7
83.4
68.5
74.6
Language
General Agent
BFCL_v4
72.9
64.14
65.16
-
76.66
72.95
67.82
-
-
-
Language
General Agent
TAU3-bench
68.3
71.1
60.43
-
47.73
56.63
64.7
-
50.67
-
Vision-Language
STEM & Puzzle
MMMU-Pro
85.9
85.72
86.24
-
-
-
82.6
83.2
83
-
Vision-Language
General VQA
MMBenchEN-DEV-v1.1
92.93
93.3
93.44
-
-
-
93.44
-
-
-
Vision-Language
Document Understanding
OCRBench
91
87.5
88
-
-
-
90.4
-
-
-
Education Benchmarks
The education evaluation covers EduBench and Pedagogy-oriented evaluation settings. The table below reports the detailed EduBench and Pedagogy Benchmark Multilingual metrics provided in the evaluation sheet.
InnoSpark3.0-397B-260712 is intended for research and application development in education-focused AI scenarios, including:
Concept explanation and step-by-step tutoring
Educational QA and homework support
Lesson planning and teaching material generation
Student-facing dialogue agents
Teacher-facing assistant workflows
General instruction following, reasoning, and agent-style tasks
Limitations
Like other large language models, InnoSpark3.0-397B-260712 may generate inaccurate, incomplete, or biased content. Outputs in educational settings should be reviewed by qualified educators when used for high-stakes learning, assessment, or student guidance. The model should not be used as the sole source for factual verification, grading decisions, psychological counseling, medical advice, legal advice, or other safety-critical decisions.
Evaluation results may vary with prompt format, decoding parameters, evaluation implementation, and data version. Users should conduct additional evaluations before deploying the model in production or classroom environments.
Main Contributions
Name
Responsibility
Personal link
Wentao Liu (刘文涛)
Training pipeline; SFT general and education data processing; education RL training