Views
No views yet
talkie-web base (same architecture as talkie-1930 but pre-trained on
web-style data). Tuned for the
mini-swe-agent interaction
format.| metric | value |
|---|---|
| pass@1 (n=3 independent eval runs) | 5.75% ± 1.04 pp |
| per-run resolved (out of 446) | 31, 23, 23 |
--model-impl transformers --max-model-len 32768 --dtype bfloat16) → mini-swe-agent (mini-extra swebench, temperature 0.7,
max_tokens=4096), graded with the swebench harness against
ricdomolm/SWE-bench_Verified-Working-Harbor.| Base model | talkie-web-13b-base (chat-token reinitialised) |
| Dataset | talkie-web-swe-100k-64k (100k SWE-smith trajectories, packed at 64k) |
| Trainer | TRL SFTTrainer via accelerate (8× A100) |
| Optimizer | adamw_torch_fused, β=(0.9, 0.95), ε=1e-8 |
| LR | 2e-5, cosine_with_min_lr, warmup 3% |
| Precision | bf16 |
| Weight decay | 0.1 |
| Max grad norm | 30 |
| Max length | 65,536 |
| Packing | bfd + padding-free |
| Loss | completion_only_loss=1 (loss only on assistant tokens) |
| Steps | 2,016 (this is ckpt-2000) |
modeling_talkie.py,
configuration_talkie.py). Load with trust_remote_code=True:1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model = AutoModelForCausalLM.from_pretrained(
4 "ricdomolm/talkie-web-coder",
5 trust_remote_code=True,
6 torch_dtype="bfloat16",
7)
8tokenizer = AutoTokenizer.from_pretrained("ricdomolm/talkie-web-coder")1vllm serve ricdomolm/talkie-web-coder \
2 --model-impl transformers --max-model-len 32768 --dtype bfloat16ricdomolm/talkie-1930-coder
— same recipe, same SFT data, but starting from a different base model.
Reaches 4.48% ± 0.69 pp on the same eval (n=5).