Gemma-4-26B-Sol-Traces-v1
Repository coding-agent model fine-tuned from unsloth/gemma-4-26B-A4B-it using LoRA on 25,000 verified deterministic reference trajectories.
Sol Traces denotes tool-use traces compiled from Hermes Agent session logs; the traces do not originate from OpenCode.
Training Details
| Parameter | Value |
|---|
| Base model | unsloth/gemma-4-26B-A4B-it (MoE, 26B total, 4 active experts) |
| Fine-tuning | LoRA (r=16, alpha=16, dropout=0) |
| Target modules | Language + attention (k/q/v/o/gate/up/down projection) |
| Dataset | 21,174 train / 1,324 val (gemma-4-native-tools format) |
| Dataset provenance | original-synthetic — 25,000 verified trajectories compiled from Hermes Agent session logs across 32,560 attempted scenarios |
| Epochs | 1 |
| Learning rate | 1e-4, cosine scheduler with 3% warmup |
| Batch size | 8 (1 × 8 gradient accumulation) |
| Max sequence | 8,192 tokens |
| Loss type | Assistant-only (tool responses excluded from loss) |
| GPU | Modal H100 80GB |
| Training time | ~2h (pilot 8m + full 1h 59m) |
| Final train loss | 0.01134 |
| Peak VRAM | 60.3 GiB / 80 GiB |
| Throughput | 3,461 tok/s |
Dataset
The training dataset consists of 25,000 executable trajectories built by a deterministic scenario generator and replayed against generated repositories. It uses 224 language/task/variant repository families with repository-family-balanced splits:
- 21,174 training records
- 1,324 validation records
- 2,502 test records (see
dataset_manifest.json)
Each trajectory is a full agent session containing:
- System instruction: Repository coding agent with tool-use guidelines
- User task: A well-scoped coding task from the deterministic fixture catalogue
- Assistant tool calls: Multi-step function-calling sequences using 5 tools:
list_files — glob-based file discovery
read_file — line-range file reading
search_code — regex code search (defined in the schema; not emitted by the v1 reference policy)
run_command — allowlisted shell execution
apply_patch — unified diff application
- Tool responses: Output, exit codes, truncation markers
- Verification: Post-task validation commands with pass/fail outcomes
Actual v1 task coverage
| Type | Records |
|---|
debugging | 5,424 |
feature | 4,709 |
refactoring | 3,582 |
testing | 3,607 |
build_config | 3,269 |
integration | 2,742 |
documentation_review | 1,667 |
Repository fixtures cover TypeScript, JavaScript, Python, shell, configuration, Go, Rust, and JVM/Java.
Data generation and verification
Sol Traces are compiled from Hermes Agent session logs produced while running deterministic, seed-based coding scenarios through a reference executor. The scenarios define repository templates, task requirements, and verification commands; accepted records retain the corresponding tool-use events and verification outcomes. Records are included only when their configured post-task validation succeeds.
The v1 reference policy is intentionally narrow: it always lists files, reads the known implementation path, runs pre-patch verification, applies the reference patch, and reruns verification. search_code is included in the schema but has no v1 calls.
Key Statistics
| Metric | Value |
|---|
| Trace source | Hermes Agent session logs (deterministic scenario generator + reference executor) |
| Attempted seeds | 32,560 |
| Accepted trajectories | 25,000 (76.8% acceptance rate) |
| Rejections | 5,872 structural duplicates + 316 verification failures |
| Provenance | original-synthetic |
| Repository families | 224 language/task/variant families across 8 fixture categories |
Files
| File | Size | Description |
|---|
gemma-4-26b-sol-traces-v1-Q4_K_M.gguf | 15.64 GiB | Quantized merged model (Q4_K_M) — ready for inference |
gemma-4-26b-sol-traces-v1-f16.gguf | 47.04 GiB | Full F16 merged model — for custom quantization |
adapter/ | — | LoRA adapter directory (safetensors, configuration, processor, and tokenizer files) |
training_stats.json | — | Full training metrics |
dataset_manifest.json | — | Accepted-record counts, split ratios, and rejection summary |
Note: The Q4_K_M file is the recommended llama.cpp deployment artifact. The published adapter is a PEFT safetensors directory, not a GGUF LoRA.
Usage (llama.cpp)
1# Q4_K_M — one file, ready to go
2llama-cli \
3 -m gemma-4-26b-sol-traces-v1-Q4_K_M.gguf \
4 -ngl 99 \
5 --prompt "List the files in the repository matching *.py"
6
7# The published adapter is safetensors, not GGUF. Load it through a PEFT/Transformers-compatible runtime.
8# For llama.cpp deployment, use the merged Q4_K_M artifact above.
Capabilities
The model excels at:
- Function calling: Selecting and populating the right tool from natural language
- Code navigation: Searching, reading, and listing files to understand codebases
- Shell execution: Running commands with proper flags and paths
- Patch application: Making small, correct code changes via unified diffs
- Deterministic verification flow: Reproducing the fixture failure, applying the reference patch, and rerunning configured checks
- Verification: Running tests and validating changes
Limitations
- Fine-tuned for repository coding agent scenarios — general chat or creative writing may not benefit
- Single-turn trajectories only — no conversational memory across separate turns
- Tool schemas are fixed to the 5 tools in the training set
- Trained on synthetic trajectories — real-world coding patterns may differ
Training Stats
1{
2 "training_loss": 0.01134,
3 "eval_loss": 0.02422,
4 "steps": 377,
5 "train_tokens": 24,704,714,
6 "peak_vram_gib": 60.3,
7 "throughput_tok_s": 3461,
8 "runtime": "1h 59m"
9}
Disclaimer
Use at your own risk. This model is fine-tuned for coding-agent scenarios. The model owner accepts no liability for any damages or losses arising from its use. Users are responsible for compliance with applicable laws and regulations.