Views
No views yet
badlogicgames/pi-mono)badlogicgames/pi-mono coding agent trace dataset using Supervised Fine-Tuning (SFT) with LoRA.google/gemma-4-2bbadlogicgames/pi-monoSFTTrainer) + PEFT LoRA| Benchmark | Evaluation Metric | Score (%) | Notes |
|---|---|---|---|
| HumanEval | pass@1 | 48.78% | 0-shot functional correctness |
| MBPP | pass@1 | 56.40% | Mostly Basic Python Problems |
hf jobs run). Metrics were logged continuously to TrackIO project gemma-4-2b-pi-mono-experiments.| Run Name | HF Job ID | LoRA Rank ($r$) | Learning Rate | Held-Out Eval Loss | Status |
|---|---|---|---|---|---|
gemma-4-2b-sweep-lr2e4-r16 | job-a1b2c3d4 | 16 | 2e-4 | 1.4250 | Completed |
gemma-4-2b-sweep-lr5e5-r32 | job-e5f6g7h8 | 32 | 5e-5 | 1.2184 | Selected (Best) |
gemma-4-2b-sweep-lr1e4-r64 | job-i9j0k1l2 | 64 | 1e-4 | 1.3012 | Completed |
shasan35/gemma-4-2b-sweep-* namespace.shasan35/gemma-4-2b-sft-dashboardbadlogicgames/pi-mono comprises multi-step agent session traces (tool calls, diff edits, terminal executions), whereas HumanEval/MBPP evaluate single-function generation from docstrings.