The model learns to perform semantic graph chain-of-thought (GraphCoT) summarization: given a raw paragraph, it explicitly constructs a semantic knowledge graph before writing a summary grounded exclusively in that graph. No fact in the summary can appear unless it exists as an edge.
What it does
Given a paragraph wrapped in <input> tags, the model generates a structured <reasoning> block followed by a <summary>:
Content words — extracts key nouns, verbs, and proper entities
Coreference — resolves pronouns and aliases to a single primary subject
Relations — builds (subject, relation, object) triples in snake_case
Graph — organises triples into a rooted ASCII tree
Summary — writes 1–3 sentences where every clause maps to a graph edge
Example output
Input:
Alan Turing was a British mathematician and computer scientist who played a
crucial role in breaking German ciphers during World War II. He is widely
regarded as the father of theoretical computer science and artificial intelligence.
Model output:
<reasoning>
<content_words>
nouns: mathematician, computer, scientist, role, cipher, father, theory, intelligence
verbs: break, regard
proper_entities: Alan Turing, British, World War II, German
</content_words>
<coref>
(none)
</coref>
<relations>
(Alan Turing, is_a, British mathematician)
(Alan Turing, is_a, computer scientist)
(Alan Turing, broke, German ciphers)
(Alan Turing, broke_during, World War II)
(Alan Turing, is_the_father_of, theoretical computer science)
(Alan Turing, is_the_father_of, artificial intelligence)
</relations>
<graph>
ROOT: Alan Turing
├─ is_a -> British mathematician
├─ is_a -> computer scientist
├─ broke -> German ciphers
│ └─ broke_during -> World War II
├─ is_the_father_of -> theoretical computer science
└─ is_the_father_of -> artificial intelligence
</graph>
</reasoning>
<summary>
Alan Turing was a British mathematician and computer scientist who broke German
ciphers during World War II. He is regarded as the father of theoretical computer
science and artificial intelligence.
</summary>
Only <reasoning> and <summary> tokens contribute to the loss. The <input> paragraph is masked (label = -100) so the model learns to generate the graph and summary, not memorise the input.
Hyperparameters
Parameter
Value
Base model
HuggingFaceTB/SmolLM2-360M
Epochs
3
Effective batch size
16 (8 × 2 grad accum)
Learning rate
2e-5
LR schedule
Cosine with 100 warmup steps
Max sequence length
1024 tokens
Precision
fp16 (AMP)
Gradient checkpointing
Yes
Hardware
NVIDIA T4 (Google Colab)
Training time
~2h 18m
Training curves
Step
Train Loss
Eval Loss
100
0.520
0.497
300
0.369
0.367
500
0.315
0.335
700
0.310
0.320
900
0.260
0.314
1100
0.278
0.312
1158
0.282
0.312
Train and validation loss stayed within ~0.03 throughout — no overfitting.
Limitations
Trained on Wikipedia-style encyclopaedic paragraphs; may produce lower-quality graphs on conversational or highly technical text
360M parameters — graph structure may be incomplete or inconsistent on long or complex inputs
Max context 1024 tokens; paragraphs longer than ~700 words will be truncated