Evaluation was performed using a custom triplet extraction benchmark with Hungarian bipartite matching alignment on 700 held-out entries.
Metrics
Metric
Score
Weight
Schema score
1.000
0.30
Entity F1
0.179
0.25
Relation accuracy
0.680
0.20
Grounding
0.969
0.15
Weight score
0.526
0.10
Triplet F1 (info only)
0.122
—
Type agreement (info only)
0.854
—
Hallucination rate
0.033
—
Composite Score: 0.6583
What Finetuning Fixed
Finetuning addressed three major failure modes of the base model.
1. Entity Normalization
Input passage:
Studies of the Cambrian period document the rapid diversification of animal life and the emergence of most major animal phyla, with some researchers proposing that a celestial body impact may have triggered the extinction events that preceded this radiation.
Base entity title extracted:
After a thorough research on the circumstantial changes and the great evolution of life in the Cambrian period
Finetuned entity title extracted:
Celestial body impact hypothesis
The finetuned model learns reusable and atomic graph nodes rather than copying passage fragments.
2. Schema Adherence
Base relations generated:
released
benefited_from
Finetuned relations generated:
based_on
used_for
applied_to
introduces
All generated relations belong to the predefined ontology.
3. Confidence Calibration
Base weights:
0.8
0.8
0.8
0.8
Finetuned weights:
0.23
0.41
0.59
0.77
The model learns meaningful confidence distributions where stronger relations receive higher scores.
Intended Use
This model is intended for:
Knowledge Graph Construction
GraphRAG pipelines
Structured Information Extraction
Entity-Relation Extraction
Automated KG population
Document-to-Graph conversion
Limitations
While the model demonstrates strong schema adherence and grounding, several limitations remain.
Shallow Entity Abstraction
The model favors concise and reusable entities but may miss deeper semantic abstractions or hierarchical entity relationships.
Limited Recall
The model prioritizes schema correctness and grounded extraction over exhaustive triplet recall. Entity F1 of 0.179 reflects strict Hungarian-matching alignment on a 20-relation ontology-constrained task; recall is intentionally traded for precision and schema adherence.
English-Centric Training
Training was primarily conducted on English Wikipedia and arXiv passages.
Ontology Constrained
Only the predefined 20 relation types are supported.
Model Size Constraints
Despite the relatively small size (0.6B parameters) and a modest training corpus (~3K examples), the model learns stable ontology-constrained extraction behavior. Larger models may achieve deeper entity understanding and broader relation coverage.