Gemma 2 2B,
instruction-tuned on top of
thesreedath/slm-gemma-2b-qa
(lineage: gemma-2-2b-it -> closed-book QA SFT -> instruction SFT). Trained on
~6.5k domain-grounded synthetic legal/financial instructions, every example
compliance- and groundedness-judged. Loss on response tokens only.