Views
No views yet
pi coding-agent
session traces from badlogicgames/pi-mono.| Field | Value |
|---|---|
| Base model | google/gemma-4-E2B-it |
| Dataset | badlogicgames/pi-mono |
| Selection metric | Lowest held-out SFT eval loss |
| Selected adapter | herooooooooo/first-hf-run-pi-mono-gemma4-e2b-adapter-lr2e-4-r16-alpha32 |
| Final model repo | https://huggingface.co/herooooooooo/first-hf-run-pi-mono-gemma4-e2b-final |
| Artifact repo | https://huggingface.co/herooooooooo/first-hf-run-pi-mono-gemma4-e2b-artifacts |
| Trackio project | first-hf-run |
| Trackio dashboard | https://herooooooooo-trackio.hf.space/?project=first-hf-run |
| Controller job id | 6a39cb35c612b71be2578806; local recovery and repair 2026-06-23T11:17:32.245324+00:00 |
| Kind | Name | Job ID | Final Stage | Link |
|---|---|---|---|---|
| sweep | lr1e-4-r8-alpha16 | 6a39cb3ec612b71be257880a | ERROR_WITH_ADAPTER_SAVED | job |
| sweep | lr2e-4-r16-alpha32 | 6a39cb3ec612b71be257880c | ERROR_WITH_ADAPTER_SAVED | job |
| sweep | lr5e-5-r16-alpha32 | 6a39cb3fc7d51fa1097d6085 | ERROR_WITH_ADAPTER_SAVED | job |
| adapter-eval | lr1e-4-r8-alpha16 | 6a3a50133d2ca349dc7bf6ac | COMPLETED | job |
| adapter-eval | lr2e-4-r16-alpha32 | 6a3a5013e902455642c9d107 | COMPLETED | job |
| adapter-eval | lr5e-5-r16-alpha32 | 6a3a5014e902455642c9d109 | COMPLETED | job |
| merge | merge-selected-adapter | 6a3a52a5f6cddbe979170025 | COMPLETED | job |
| eval | humaneval | 6a3a5750f6cddbe97917004e | COMPLETED | job |
| eval | mbpp | 6a3a5750f6cddbe979170050 | COMPLETED | job |
| Run | Adapter repo | Held-out eval loss | Job id reported by child |
|---|---|---|---|
lr1e-4-r8-alpha16 | herooooooooo/first-hf-run-pi-mono-gemma4-e2b-adapter-lr1e-4-r8-alpha16 | 2.2324344871670063 | 6a3a50133d2ca349dc7bf6ac |
lr2e-4-r16-alpha32 | herooooooooo/first-hf-run-pi-mono-gemma4-e2b-adapter-lr2e-4-r16-alpha32 | 2.0485327514494878 | 6a3a5013e902455642c9d107 |
lr5e-5-r16-alpha32 | herooooooooo/first-hf-run-pi-mono-gemma4-e2b-adapter-lr5e-5-r16-alpha32 | 2.350842223989554 | 6a3a5014e902455642c9d109 |
| Benchmark | Accuracy | StdErr | Completed | Return code |
|---|---|---|---|---|
| humaneval | 0.0 | 0.0 | 164/164 | 0 |
| mbpp | 0.0 | 0.0 | 1285/1285 | 0 |
evals/.-M chat_template=... because the stock Inspect HF provider handed ChatMessage objects to a dict-oriented tokenizer template.max_connections=8 on l4x1 to complete the full MBPP run within the HF Job timeout.eval_results.json.accuracy and stderr; it does not publish separate pass@k columns.l4x1 with max_connections=8, not vLLM throughput.local inside the HF Job, not Docker; this is less isolated than leaderboard-grade Docker execution.pi-share-hf, including deterministic
redaction, deny-pattern filtering, TruffleHog scanning, and LLM review before
upload. This run consumes the published redacted dataset only.run_scripts/. The article-style write-up is in ARTICLE.md.