This is the canonical reproducible LoRA adapter for the summa-moral-graph Christian virtue SFT baseline. It fine-tunes Qwen/Qwen2.5-1.5B-Instruct toward Aquinas-grounded Christian virtue reasoning while preserving explicit passage-level traceability to reviewed evidence.
Abstract
The purpose of this model is not to produce generic theological chat or to memorize citation strings. The goal is to show, in a compact and reproducible public baseline, that reviewed Summa Moral Graph supervision can measurably move a general model toward Aquinas's virtue categories, evidence-bounded answers, and citation-aware outputs.
Related Projects
Summa Moral Graph
I structured the moral reasoning in Thomas Aquinas's Summa Theologiae as a knowledge graph: browsable, downloadable, and open source.
Built on that graph, I fine-tuned a small 1.5B Qwen model with supervised fine-tuning on the reviewed dataset, kept the workflow open source, and exposed a public chat demo.
LoRA on Apple Silicon mps, float16, no quantization
Reviewed source annotations
555
Total SFT examples
1883
Train / val / test
1475 / 175 / 233
Canonical run id
20260421_134712
Git commit
40c724d0aaab5cdedc25110a1b4545157e9dcea3
Held-out exact citation
36.5%
Strongest task slice
Virtue concept explanation at 65.6%
Strongest tract slice
Justice core at 50.0%
Artifact Status
The public GitHub release keeps the earlier distribution tag christian-virtue-qwen2.5-1.5b-local-baseline-20260418_193038 for continuity, but the authoritative benchmark numbers in this package and curated report come from the corrected run 20260421_134712.
Treat the curated report and local package manifest as the canonical evaluation surface for the current repo numbers.
subset_summary.json records the exact balanced (task_type, tract) composition of the local training and eval subsets used for this run.
Why This Adapter Exists
Train an Aquinas-grounded Christian virtue assistant rather than a generic theology bot.
Keep the supervision evidence-first: reviewed doctrinal annotations only, joined back to stable passage ids.
Demonstrate a small public baseline that others can inspect, reproduce, and adapt to their own models before scaling up to larger runs.
Public Benchmark Highlights
Highlight
Base
Adapter
Delta
Held-out benchmark exact citation
0.0%
36.5%
36.5%
Virtue concept explanation
0.0%
65.6%
65.6%
Reviewed relation explanation
0.0%
62.7%
62.7%
Justice core tract
0.0%
50.0%
50.0%
Executive Readout
Held-out benchmark exact citation reaches 36.5% over 233 prompts.
The clearest public win is Virtue concept explanation: 65.6% exact over 32 held-out prompts.
Second strongest task slice: Reviewed relation explanation at 62.7% exact over 67 prompts.
Strongest tract slice: Justice core at 50.0% exact over 42 prompts.
This published run uses a deliberately small 1.5B local demo model, so the result should be read as proof that the pipeline works rather than as the ceiling for final quality.
This package intentionally foregrounds the strongest virtue-aligned slices; the full held-out matrix remains in the published report.
Full task/tract breakdowns and the qualitative goal-demo panel live in the published report.
Held-out benchmark comparison
Figure. Held-out base-vs-adapter comparison from the canonical local local-baseline run. The key claim is straightforward: even a small reproducible demo baseline moves model behavior in the right direction, which makes this a credible public SFT template rather than only a code release.
Dataset And Evidence Policy
Training export: data/processed/sft/exports/christian_virtue_v1