SmartSight Coach LoRA fine-tunes
LoRA fine-tunes for the
SmartSight AI coach — an on-device fitness and nutrition coach that
runs entirely on the phone, no account and no server round-trip. Converted to
.litertlm for
Android/iOS inference. →
smartsight.app/ai-coach
Built and maintained by
Niclas Bade.
Base model changed at v42 (2026-08-12). v27–v38 were fine-tuned from google/gemma-4-E2B-it.
Everything from v42 on is fine-tuned from Google's quantization-aware-trained checkpoint
google/gemma-4-E2B-it-qat-q4_0-unquantized, whose weights are conditioned to survive 4-bit
rounding. The frontmatter above reflects the CURRENT parent, so the model tree matches what actually
shipped. This repo keeps the current best checkpoints plus earlier superseded milestones for
history — not every candidate tested each round. Rounds that lose are documented internally with
their real failure cases; they are never uploaded here.
More coming. The vision tower is being trained surface by surface — body fat, meal photos and
nutrition-label reading — each measured against held-out data before anything ships. Progress and
honest limitations for every round are written up below and at
smartsight.app/ai-coach.
Reverted promotion — v164ctl456 (2026-09-04, same day)
v164ctl456 was promoted on vision+text totals (189.0 vs 187.5) and REVERTED hours later: the
project's scoring matrix includes the nutrition-label surface (labelReal32, 32 real phone/shelf
captures, estimator = scan accepted by the app AND kcal correct), and on the full matrix the
standing champion wins — v158 totals 196.5 (label 9/32) vs v164ctl456 195.0 (label 6/32). The
challenger's +1 vision and +0.5 text were bought with -3 label reads and one background-probe
fabrication. Nothing about v158 changed; the error was promoting on an incomplete total, and this
entry stays on the card so the history is honest.
smart-coach-vision.litertlm — CURRENT CHAMPION (v158)
Promoted 2026-09-02. sha256 3a9f9798b7fe2bfa, 2,795,655,584 bytes.
Lineage: v27 → v31 → v34 → v38 → v42 → v44 → v46 → v47 → v48 → v51 → v54 → v55 → v69 → v85 →
v127 → v134 → v137 → v138 → v139 → v158.
It stopped answering different bodies with the same number. v139 closed the lean end and lifted
ranking correlation; its remaining defect was COLLAPSE — a single emitted value absorbing an
11.9-point range of true body fat. v158 nearly halves that, improves accuracy and ranking, and takes
coaching text to near ceiling.
Both models re-measured head to head on the same frozen 49-photograph surface, same prompt, same
backend, matched row by row.
| surface | v139 | v158 |
|---|
| within 2 points, 49 real photographs | 30/49 | 32/49 |
| mean error, 49 real photographs | 2.11 | 1.76 pp |
| ranking correlation, 49 photographs | +0.939 | +0.958 |
| rank correlation (Spearman) | +0.926 | +0.941 |
| widest attractor span | 11.9 | 6.1 pp |
| regression slope (1.0 = tracks truth) | 0.97 | 0.98 |
| bias (negative = reads low) | −0.78 | −0.87 pp |
| declines non-body photographs | 30/30 | 30/30 |
| coaching text, seed 42 / seed 43 | 116 / 116 | 125 / 126 of 128 |
| distinct values emitted | 16 | 12 |
What the attractor span measures, and why it is the headline. For every value the model emits
three or more times, it is the range of true body fat absorbed by that one answer; the figure
reported is the widest. v139 had one number standing in for bodies 11.9 points apart — fluent,
confident, and unable to separate them. At 6.1 that collapse is roughly halved. Within-2pp alone
would not show this: a model can score well on that metric while bucketing, which is why slope,
ranking correlation and span are all reported beside it.
Ranking held while accuracy improved. Correlation between true and predicted body fat rises
+0.939 → +0.958, and Spearman +0.926 → +0.941, so the ordering of bodies did not degrade to buy the
accuracy gain — the failure mode that produced this project's earlier false wins.
Coaching text is where the promotion was actually won. 116 → 125/126 of 128 across two seeds,
over four suites: relative-date reads, logged-session reads, absent-day handling, and out-of-window
questions. The 0.35 pp accuracy gain is INSIDE this project's measured two-sigma noise floor of
0.902 pp, so on body fat alone the two models are not separable. The promotion rests on the totals
and chiefly on text — stated plainly, because the alternative is presenting a draw as a win.
Non-body refusals held at 30/30. That slice is a REGRESSION GUARD, not a discriminator: gains in
body-fat accuracy have historically been paid for in refusals, so it is measured every round and
carried in the total. Holding it flat is the result being reported, not a null.
Champion history — what each round actually fixed
Every entry here was a promoted champion, verified against the previous one on a held-out suite
with a full read of every answer. Rejected rounds (v28-v30, v32, v33, v35-v37, v39, v40, v41) are
not listed — they are documented internally with their real failure cases.
| Round | Beat | What it fixed | Base |
|---|
| v27 | baseline | first winning fine-tune — shorter, equally accurate general advice | gemma-4-E2B-it |
| v31 | v27 | exercise-form answers reached the untouched baseline's zero-error record (v27 was wrong on 4/5 direct form questions); no more garbled or duplicated plan output. The unlock was the export recipe, not data — three consecutive data-only rounds failed first, then Hadamard-rotation int4 fixed it with zero retraining | gemma-4-E2B-it |
| v31 export update | — | same weights, AlgorithmName.HADAMARD_ROTATION custom op instead of the decomposed variant: 88% faster decode, ~85 MB smaller, quality statistically unchanged (5.8% vs 6.4% defect rate over 312 generations) | — |
| v34 | v31 | dietary-rule compliance (a "vegetarian, no nuts" request had served peanut butter); first correct full examples for face pull and plank; squat-vs-hip-thrust contrast cases | gemma-4-E2B-it |
| v38 | v34 | AI challenges: unrealistic targets for low-frequency goals (150 progress photos in six weeks from zero) and non-English requests answering in English. Won over its own alternate seed on a Danish grammar defect — plural "jeres/jer" where the app addresses one person | gemma-4-E2B-it |
| v42 | v38 | see above — QAT base, blockwise export, and the fenced-JSON data fix. Won over its own alternate seed, which produced a confident wrong number (subtracting a 78 kg bodyweight goal from a 106.7 kg squat 1RM) | gemma-4-E2B-it-qat-q4_0-unquantized |
| v44 | v42 | Danish macro vocabulary and challenge titles. Measured cause: of 7,288 corpus rows, 13 contained any Danish/Norwegian, zero taught "kulhydrat", and ten demonstrated exactly the wrong behaviour — English macro words inside Danish prose. Localising those ten rows fixed it; /nutrition emission went 5/7 → 7/7 | gemma-4-E2B-it-qat-q4_0-unquantized |
| v46 | v44 | avoid-list violations eliminated (0/36 vs 5/36) by matching the corpus to the prompt format production actually sends | gemma-4-E2B-it-qat-q4_0-unquantized |
| v47 | v46 | no retraining — the vision export ran at 140 soft tokens instead of the 280 the checkpoint specifies, so every shipped build saw 48.9% of the pixel area. The asymmetry proved the mechanism: re-exporting at 280 moved body-fat rank correlation in men +0.492 → +0.704 while women were unchanged, because what separates adjacent bands in men is fine texture only a few pixels tall and downscaling is a low-pass filter | gemma-4-E2B-it-qat-q4_0-unquantized |
| v48 | v47 | conversational follow-ups: the coach dropped the app's own numbers as soon as a user pushed back. Failures 16/24 → 8/24 | gemma-4-E2B-it-qat-q4_0-unquantized |
| v51 | v48 | stopped abandoning the app's numbers on a user contradiction, and stopped inventing plausible detail it could not see ("right in the middle of where it usually sits" → "I can't tell what your usual is") | gemma-4-E2B-it-qat-q4_0-unquantized |
| v54 | v51 | read the wrong cell of the 7-day history line: asked "what did I train yesterday?" it answered about TODAY, 18 times out of 32 | gemma-4-E2B-it-qat-q4_0-unquantized |
| v55 | v54 | the body-fat read was a CONSTANT — 18% for all 42 test photos, and also for a grey rectangle, for pure noise, and for no image at all. Cause was the training data, not the vision tower: a text-only LoRA destroys vision the base model already has (sensitivity fell 0.1164 → 0.0405 across v51→v54 with ZERO vision tensors in the adapter — the language layers consume the image tokens). Mixing ~17% of the vision surface's rows back into the text corpus restored it: MAE 9.13 → 3.42, sensitivity 0.0405 → 0.2183, grey rectangle finally separating from real photos. ⚠️ It also lost the base model's striation detection while gaining the percentage read — capabilities traded, not accumulated, which is why every surface is probed at promotion now | gemma-4-E2B-it-qat-q4_0-unquantized |
| v69 | v55 | relative-date reads ("what did I train yesterday") 24/32 → 32/32. v55 had a consistent off-by-one — it treated the last day in the 7-day line as yesterday instead of today, then described the wrong row accurately: fluent, confident, wrong. Six earlier rounds tried to train it away and every one made it worse. What worked was teaching the model to STATE THE ANCHOR BEFORE USING IT ("the line ends with Sun, so today is Sun; one back is Sat"), turning an indexing problem into two lookups it could already do. Body fat on 37 unseen ordinary phone photos (bathroom mirrors, kitchens, garages — not studio imagery) 2.88 → 1.88 pp, and the most common single answer fell 43% → 19% of photos. ⚠️ It also introduced a regression, fixed by addition not retreat: 100% of the new rows ask the model to locate a day BEHIND today, so it over-generalised to "behind today = not current" and started saying a logged past session doesn't count (32/32 → 27/32, identical on all three seeds) | gemma-4-E2B-it-qat-q4_0-unquantized |
| v85 | v69 | body-fat reads on real photographs stopped compressing the range: slope 0.815 → 1.043, mean error 1.99 → 1.69 pp, and the high end went from reading 24% as 21 to reading it as 24. Coach text and follow-ups level on two seeds. ⚠️ the lean end did not move at all — 5.9%, 8.8% and 10.0% all still read 8, exactly as in v69 | gemma-4-E2B-it-qat-q4_0-unquantized |
| v127 | v85 | gained a category for "this is not a person", the open defect since v55. A plain grey rectangle went from a confident "12%" to "I cannot see a person in this image"; real bodyless photographs — walls, sofas, pets, food — went from 0/30 declined to 14/30. Also fixed background-swap instability: the same body could be moved 10 points by changing the background in v85, worst case 4.0 in v127. The measured cause is the training set, not the objective: arms trained on synthetic greys alone scored 0/30 on real photographs, identical to v85, while arms given 15 real bodyless photographs fabricated 31.6 pp less (seed-paired, bar set in advance at 22 pp). ⚠️ Still fabricates on 16 of 30 real bodyless photographs, and the promoted seed is not the strongest of the three on this surface | gemma-4-E2B-it-qat-q4_0-unquantized |
| v134 | v127 | Shipped as the downloadable weights on 2026-08-29 but NEVER GIVEN A CARD ROW UNTIL NOW - this entry is added retrospectively so the lineage is not missing a link. Measured on the current evaluation set at promotion time of its successor: composite error 2.58 pp across three photo sets, bodyless fabrication 12 of 58, background-robustness 46 of 58. Its gain over v127 is NOT restated here because v127 was never re-measured on this evaluation set, and a number that was not measured does not go on a public card. | gemma-4-E2B-it-qat-q4_0-unquantized |
| v137 | v134 | mean error across three photo sets 2.58 -> 2.08 pp and invented readings on non-body photos 12/58 -> 3/58. Cause was the training labels, not the recipe: 163 photographs re-judged by a single stronger judge with anchors packed inside the band being judged, plus new real people. The round's own experiment (real-photo share) was REJECTED - the gain came from label quality | gemma-4-E2B-it-qat-q4_0-unquantized |
| v138 | v137 | invented readings on non-body photos 3/58 -> 2/58, and background instability cut threefold - the same body on a different background moved 10 points worst-case in v137, 4 in v138. WARNING it REGRESSED on body-fat accuracy: within 2 points on 64% -> 54% of photographs, and the 47-photo set went 2.12 -> 2.99 pp reading about 1.9 low. Promoted on a saturated matrix by a one-image margin | gemma-4-E2B-it-qat-q4_0-unquantized |
| v139 | v138 | half points, and the lean end — the open defect since v69. v138 answers the 7.5% DEXA reference photo as 6 or 10 and cannot say 7.5 under any prompt, because the capability was never trained; v139 answers 7.5 exactly on 4 of 9 encodings, mean error on that photo 2.39 → 1.22 pp and bias +2.06 → +0.78. Also declines 58/58 non-body photographs against v138's 56/58, and takes the counted vision total 131/133 → 133/133 with no surface regressing. Ranking correlation over 12 distinct men +0.688 → +0.806 with the regression slope 0.777 → 0.930. ⚠️ its training data contradicts itself — 889 rows say "ONE WHOLE NUMBER" and answer "7.5%" — so it is unusually prompt-sensitive and needs the prompt printed above | gemma-4-E2B-it-qat-q4_0-unquantized |
| v158 | v139 | stopped collapsing distinct bodies onto one answer — widest attractor span 11.9 → 6.1 pp on the 49-photograph surface, mean error 2.11 → 1.76 pp, within-2pp 30/49 → 32/49, ranking correlation +0.939 → +0.958 (Spearman +0.926 → +0.941), and coaching text 116 → 125/126 of 128 across two seeds with non-body refusals held at 30/30. Won on the TOTALS and chiefly on text: the 0.35 pp accuracy gain is inside the measured 0.902 pp two-sigma floor. WARNING its own tested variable FAILED its precommit (−0.160 against a 0.553 bar), so this was a better draw and not a demonstrated cause; bias also went −0.78 → −0.87 pp and distinct values emitted fell 16 → 12 | gemma-4-E2B-it-qat-q4_0-unquantized |
| v164ctl456 (REVERTED same day) | v158s123 | widest attractor span 6.1 to 5.6 and measured within-2pp 32/49 to 34/49 at slope 0.996 with MAE 1.76 to 1.66; coaching text 125.5 to 126 of 128. Promoted from a CONTROL arm: the gain is a seed draw, not the rounds variable (fig-coverage precommit FAILED and is reported as such). Known note: one background-probe fabrication (a sofa read as 18.5 percent), 30/30 to 29/30. | gemma-4-E2B-it-qat-q4_0-unquantized |
Recipe constant across every round since v22c: LoRA r=16 / alpha=32 / dropout=0.1, 2 epochs,
lr 6e-5, NEFTune noise_alpha=5, completion-only loss masking (train_on_responses_only),
per-task validation shards, two seeds per round with the winner chosen on behaviour rather than
validation loss. Only the training data, the base, and the export recipe have moved.
Honest current weaknesses
Documented rather than hidden, because they are the targets for the next rounds:
-
Body fat at the lean end: LARGELY FIXED IN v139, and the fix was half points. Through v138
this read as a bucket — DEXA-measured 5.9%, 8.8% and 10.0% photographs all returned 8, three
truths and one answer, unchanged from v69 to v85. The cause turned out to be partly the OUTPUT
ALPHABET: a model that can only answer in whole numbers cannot express 7.5, and prompting it to
do so changes nothing because the capability was never trained. v139 was trained with half-point
targets and answers 7.5 exactly on 4 of 9 encodings of the DEXA reference photograph, cutting mean
error there 2.39 → 1.22 pp. Remaining gap: this is demonstrated on ONE subject at nine
encodings. More DEXA-measured lean subjects is still the lever for proving it generalises.
-
Images that are not bodies: LARGELY FIXED IN v127, not yet solved. This was listed here
as unfixable-by-prompting through v85, and it was: a variant instructing the model to reply
NOT_A_BODY was ignored, returning byte-identical output. What fixed it was training data —
specifically real bodyless photographs, not synthetic ones. A plain grey rectangle now reads
"I cannot see a person in this image" rather than "12%". Remaining gap: v127 still returns a
percentage for 16 of 30 real bodyless photographs (walls, sofas, pets, plates of food), so the
behaviour is much better and still not dependable. More real bodyless photographs is the known
lever; synthetic ones were measured to buy almost nothing.
v139 update: it now declines all 58 of the bodyless probe set, answering "no person visible"
every time, against v138's 56/58. Two cautions: the probe set is largely SYNTHETIC (grey ramps and
noise), which is exactly the imagery measured to prove little, and the behaviour depends on the
prompt carrying an explicit refusal clause. Real bodyless photographs remain the outstanding fix.
-
The training prompt contradicts the training targets (v139). Half-point targets were added
without updating the instruction the model is trained against, so 889 of 4,494 rows pair "ONE
WHOLE NUMBER" with an answer of "7.5%". One row in five teaches the model to disobey its own
instruction, which is why v139's readings move by more than a point on small prompt edits. The
next round tests the repair directly, treated against its own control.
-
Meal-plan portions are sized by habit, not arithmetic. The model anchors on ~100 g
portions rather than solving the prompt's stated calorie budget. Mostly invisible in the app,
which rescales a day to its target (MealPlanScaler, clamped 0.6x–2.0x) — but one case per
model is still off after that clamp, and the underlying arithmetic is unsolved.
-
Avoid-lists are not always respected — foods explicitly banned for variety reappear
(2/29 held-out cases on v44, v42 and v38 alike; this is the single defect class that has
survived every round so far).
-
Nordic follow-ups about an UNLOGGED past activity regressed in v44. Asked in Norwegian
about a hike "yesterday" that is not in the app, v42 correctly said the gap was only a logging
omission; v44 answers as though advising about today. 26 Nordic follow-up rows were added to
fix exactly this and did not move it — the cause is not simply coverage.
-
Occasional Danish/Norwegian word-formation slips in v44 — "vækning" (not a word) and the
hybrid "fat-mål". The macro vocabulary itself is now correct, but morphology is not reliable.
-
The /nutrition command line is improvised. The corpus contains ZERO examples of it; the
format lives only in the prompt. v44 emits the line on 7/7 applicable cases (v42 omitted it
entirely on 2/7) but drops the trailing fat value on 3/7, so the app applies three of the
four fields. Harmless by design — NutritionCommandParser accepts any subset — but the fat
target silently stays stale.
-
Within-day duplicate or misplaced exercises still appear occasionally in generated
routines (e.g. a chest isolation movement landing in a Legs day). Lower real-world impact
than it sounds, since the app's own parser drops duplicate slots.
-
Confabulation-despite-coverage on a few specific lifts (squat/hip-thrust blending,
farmer's carry) that does not reliably respond to more training data.
-
Superset pairing — avoiding two competing-muscle compounds back to back — is imperfect.
-
Tight macro-budget precision (±8–10%) is unsolved across every model tested, including
the untouched baseline.
Evaluation discipline
Every promote/reject decision reads every row of the eval suite, never a sample — an early
round reported a verdict from ~12 of 52 rows and missed hard timeout loops entirely. Prose
quality is judged as its own axis alongside structural correctness, because a structurally
clean answer can still be flat, templated, or subtly wrong in a second language. Candidates are
always compared against both the previous champion and the untouched baseline.
Files
Which file the app fetches
Two files are published. The app decides which one to fetch based on its build version; there is
nothing to select and nothing to configure.
| file | who gets it | size |
|---|
smart-coach-vision.litertlm | ⭐ current champion — text coaching AND vision (body-fat photos, meal photos, nutrition labels) in one model. Fetched by the next app release. | 2.80 GB |
smart-coach.litertlm | text-only, fetched by app builds already installed. Unchanged on purpose. | 2.63 GB |
Both carry the same v46 coach weights (v47 changes only the vision export, not the weights), so an older build is not stuck on an older coach. They
exist side by side because installed builds also fetch a separate 3.66 GB vision model — handing
them the larger consolidated file would raise their total download on devices already tight on
RAM. See the current-champion section above for what the consolidation actually changes.
smart-coach-vision.litertlm — ⭐ current champion, fetched by the next app release. The v46 coach WITH a working vision
tower: one model for text coaching, body-fat photos, meal photos and nutrition labels (2.80 GB).
The next app release points both the coach and vision URLs here and stops downloading the
separate 3.66 GB vision model.
smart-coach.litertlm — STABLE, text-only (2.63 GB). The file app builds already in the wild
download at runtime on Android and iOS, and the one they keep using. Deliberately left unchanged:
those builds also fetch the separate vision model, so giving them the larger consolidated file
would raise their total download on devices that are already RAM-constrained. Same v46 coach
weights, so nobody is stuck on an older coach for stability's sake.
coach.litertlm — the original working int8 build from the v1 era (2.59 GB), kept as the
historical starting point.
coach-finetuned-int4.litertlm — the first int4 conversion (2.56 GB), kept because it is the
artifact that exposed the LiteRT-LM chat-template .get() incompatibility.
model.safetensors + config.json + tokenizer files — HF-format artifacts for reference.