Views
No views yet
LiquidAI/LFM2.5-350M, trained with TRL SFT on tau/commonsense_qa train examples.Answer:. The completion target is the single correct letter. Evaluation scores each candidate letter by summed log probability and chooses the highest-scoring answer.1221 examples).| Model | Forced-choice validation accuracy | Correct / Total | Job |
|---|---|---|---|
LiquidAI/LFM2.5-350M | 0.20393120393120392 | 249 / 1221 | 6a358b823093dba73ce2a757 |
LiquidAI/LFM2.5-350M + this adapter | 0.25225225225225223 | 308 / 1221 | 6a358b823093dba73ce2a757 |
+0.048321048321048304.0.2571662571662572 after SFT, but failed before artifact upload when post-train Trackio logging raised. The released checkpoint above is from the completed upload job.SFTTrainer with completion-only loss and PEFT LoRA.LiquidAI/LFM2.5-350M at commit 7728373d9f752dc3669ee3bf70786aef397874bb.tau/commonsense_qa, train split capped at 2000 examples.tau/commonsense_qa validation split, all 1221 examples.1e-4, best checkpoint by validation eval_loss.6a358b823093dba73ce2a757.burtenshaw/lfm25-commonsenseqa-sft-trackio.benchmark_before.jsonbenchmark_after.jsonbenchmark_summary.json