Views
No views yet
v0.1: research artifact, not a production model.This is the first iteration in a four-part methodology series. Form (5-7-5 syllable compliance) did NOT improve to a statistically significant degree under SFT (McNemar's exact test, p=0.167, n=100, n_discordant=19). Content metrics improved substantially. v0.1 is published as a reproducibility anchor for the lgtm-575 blog series, not as a tool for production code review.The primary deliverable of this project is the eval harness, not these weights. v1.0 (reasoning-token architecture) is in progress.
| Metric | Base floor | v0.1 | Golden number | Verdict |
|---|---|---|---|---|
| #1 Form (valid 5-7-5) | 14% (14/100) | 7% (7/100) | >= 45% | Miss |
| #3 Relevance (real-pair mean) | 0.321 | 0.394 | >= 0.37 | Hit |
| #2 Category (4-class macro-F1) | 0.319 | 0.458 | hold gap (>= +0.139) | Hit |