Gemma-4-E2B-Sol-Traces-v1
Repository coding-agent model fine-tuned from unsloth/gemma-4-E2B-it using LoRA on 25,000 verified deterministic reference trajectories.
The E2B variant is the smallest model in this four-run family. End-to-end tool-use benchmark comparisons are not yet published.
Sol Traces denotes tool-use traces compiled from Hermes Agent session logs; the traces do not originate from OpenCode.
Training Details
| Parameter | Value |
|---|
| Base model | unsloth/gemma-4-E2B-it (MoE, 2 active experts) |
| Fine-tuning | LoRA (r=16, alpha=16, dropout=0) |
| Target modules | Language + attention (k/q/v/o/gate/up/down projection) |
| Dataset | 21,174 train / 1,324 val (gemma-4-native-tools format) |
| Dataset provenance | original-synthetic — 25,000 verified trajectories compiled from Hermes Agent session logs across 32,560 attempted scenarios |
| Epochs | 1 |
| Learning rate | 1e-4, cosine scheduler with 3% warmup |
| Batch size | 8 (2 × 4 gradient accumulation) |
| Max sequence | 8,192 tokens |
| Loss type | Assistant-only (tool responses excluded from loss) |
| GPU | Modal H100 80GB |
| Training time | ~45 min (pilot 3min + full 42min) |
| Final train loss | 0.0229 |
| Validation loss | 0.0248 |
| Peak VRAM | 33.7 GiB / 80 GiB |
| Throughput | 4,587 tok/s |
These are reported run metrics; the canonical training_stats.json artifact is not currently published for E2B.
Dataset
The training dataset consists of 25,000 executable trajectories built by a deterministic scenario generator and replayed against generated repositories. It uses 224 language/task/variant repository families with repository-family-balanced splits:
- 21,174 training records
- 1,324 validation records
- 2,502 test records (see
dataset_manifest.json)
Each trajectory is a full agent session containing:
- System instruction: Repository coding agent with tool-use guidelines
- User task: A well-scoped coding task from the deterministic fixture catalogue
- Assistant tool calls: Multi-step function-calling sequences using 5 tools:
list_files — glob-based file discovery
read_file — line-range file reading
search_code — regex code search (defined in the schema; not emitted by the v1 reference policy)
run_command — allowlisted shell execution
apply_patch — unified diff application
- Tool responses: Output, exit codes, truncation markers
- Verification: Post-task validation commands with pass/fail outcomes
Actual v1 task coverage
| Type | Records |
|---|
debugging | 5,424 |
feature | 4,709 |
refactoring | 3,582 |
testing | 3,607 |
build_config | 3,269 |
integration | 2,742 |
documentation_review | 1,667 |
Repository fixtures cover TypeScript, JavaScript, Python, shell, configuration, Go, Rust, and JVM/Java.
Data generation and verification
Sol Traces are compiled from Hermes Agent session logs produced while running deterministic, seed-based coding scenarios through a reference executor. The scenarios define repository templates, task requirements, and verification commands; accepted records retain the corresponding tool-use events and verification outcomes. Records are included only when their configured post-task validation succeeds.
The v1 reference policy is intentionally narrow: it always lists files, reads the known implementation path, runs pre-patch verification, applies the reference patch, and reruns verification. search_code is included in the schema but has no v1 calls.
Key Statistics
| Metric | Value |
|---|
| Trace source | Hermes Agent session logs (deterministic scenario generator + reference executor) |
| Attempted seeds | 32,560 |
| Accepted trajectories | 25,000 (76.8% acceptance rate) |
| Rejections | 5,872 structural duplicates + 316 verification failures |
| Provenance | original-synthetic |
| Repository families | 224 language/task/variant families across 8 fixture categories |
Files
| File | Size | Description |
|---|
gemma-4-e2b-sol-traces-v1-Q4_K_M.gguf | 3.18 GiB | Quantized merged model (Q4_K_M) — recommended for deployment |
gemma-4-e2b-sol-traces-v1-f16.gguf | 8.64 GiB | Full F16 merged model — for custom quantization |
dataset_manifest.json | — | Accepted-record counts, split ratios, and rejection summary |
Note: The Q4_K_M file is the recommended deployment format. The F16 is provided for downstream quantization experiments.
Usage (llama.cpp)
1# Q4_K_M — one file, ready to go
2llama-cli \
3 -m gemma-4-e2b-sol-traces-v1-Q4_K_M.gguf \
4 -ngl 99 \
5 --prompt "List the files in the repository matching *.py"
6
7# With conversation template
8llama-cli \
9 -m gemma-4-e2b-sol-traces-v1-Q4_K_M.gguf \
10 -ngl 99 \
11 --temp 0.2 \
12 --chat-template gemma \
13 -p "Search the codebase for any TODO comments"
Capabilities
The model excels at:
- Function calling: Selecting and populating the right tool from natural language
- Code navigation: Searching, reading, and listing files to understand codebases
- Shell execution: Running commands with proper flags and paths
- Patch application: Making small, correct code changes via unified diffs
- Deterministic verification flow: Reproducing the fixture failure, applying the reference patch, and rerunning configured checks
- Verification: Running tests and validating changes
Comparison with Other Sol-Traces Models
| Model | Active Params | Q4 Size | Training Loss | Speed | Best For |
|---|
| E2B (this) | ~5B | 3.2 GB | 0.0229 | Fastest | Edge, CPU+GPU hybrid, low-resource |
| 12B Unified | 12B | 6.8 GB | 0.0800 | Fast | Balanced performance |
| E4B | ~8B | 4.9 GB | 0.0096 | Fast | Best quality-size trade-off |
| 26B-A4B | ~8B* | 15.6 GB | 0.0113 | Moderate | Maximum capability |
*E4B and 26B-A4B both activate 4 experts but have different base architectures (dedicated encoder vs unified).
Limitations
- Fine-tuned for repository coding agent scenarios — general chat or creative writing may not benefit
- Single-turn trajectories only — no conversational memory across separate turns
- Tool schemas are fixed to the 5 tools in the training set
- Trained on synthetic trajectories — real-world coding patterns may differ
Training Stats
1{
2 "training_loss": 0.0229,
3 "eval_loss": 0.0248,
4 "steps": 377,
5 "train_tokens": 24,704,714,
6 "peak_vram_gib": 33.7,
7 "throughput_tok_s": 4587,
8 "runtime": "44m 46s"
9}
Disclaimer
Use at your own risk. This model is fine-tuned for coding-agent scenarios. The model owner accepts no liability for any damages or losses arising from its use. Users are responsible for compliance with applicable laws and regulations.