Views
No views yet
step-100
checkpoint of the formatrefresh_continue200_len192_layers0_5_qo_r8 tau2
continuation run in the online-memory checkpoint repo:xiaol/gemma-4-e4B-hybrid-rnn-mem-rwkv-fable5-gpt5.5-v1fable5_gpt55_data entry in the metadata is a planned upgrade note, not the
data used for this accepted checkpoint.1https://github.com/xiaol/llama.cpp-online-memory
2commit: 85da0c63b Add Gemma4 RWKV-MS GGUF sidecar runtime
3base upstream: ggml-org/llama.cpp 1ec44d1| File | Size | SHA-256 |
|---|---|---|
gemma-4-E4B-it-Q8_0.gguf | 8,031,240,160 bytes | fb8f0c032de00b18c710824af3c7e5777c71e5fb60b13f13575f0a9e92ddecd0 |
mmproj-gemma-4-E4B-it-Q8_0.gguf | 559,874,528 bytes | 51d4b7fd825e4569f746b200fccc5332bf914e8ef7cbe447272ce4fec6df3db6 |
gemma-4-E4B-it-rwkv-ms-memory.gguf | 1,663,840 bytes | 0c646a776b5b12c9d3657ffd2e5e581be1eb46e858f1f404afeaa7077c02974e |
1/path/to/llama-server \
2 -m gemma-4-E4B-it-Q8_0.gguf \
3 --mmproj mmproj-gemma-4-E4B-it-Q8_0.gguf \
4 --alias gemma-4-e4b-it-q8 \
5 --host 127.0.0.1 \
6 --port 8080 \
7 --ctx-size 8192 \
8 --n-gpu-layers 999 \
9 --jinja \
10 --reasoning off1/path/to/llama-server \
2 -m gemma-4-E4B-it-Q8_0.gguf \
3 --alias gemma-4-e4b-it-rwkv-ms-q8 \
4 --host 127.0.0.1 \
5 --port 18083 \
6 --ctx-size 8192 \
7 --batch-size 2 \
8 --ubatch-size 1 \
9 --n-gpu-layers 999 \
10 --parallel 1 \
11 --jinja \
12 --reasoning off \
13 --rwkv-ms-sidecar gemma-4-E4B-it-rwkv-ms-memory.gguf \
14 --no-cont-batching \
15 --no-context-shift \
16 --no-cache-prompt \
17 --no-cache-idle-slots \
18 --cache-ram 0 \
19 --ctx-checkpoints 0 \
20 --slot-save-path ./slots0-5.q and o delta hooks only.num_state_heads=1.rwkv_ms_num_states=4, chunk size 1024.--parallel 1 and --ubatch-size 1.0 exact-prefix
restore checks.1ok: true
2base_gguf_sha256: fb8f0c032de00b18c710824af3c7e5777c71e5fb60b13f13575f0a9e92ddecd0
3sidecar_base_gguf_sha256: fb8f0c032de00b18c710824af3c7e5777c71e5fb60b13f13575f0a9e92ddecd01temperature=0, top_k=1, top_p=1, seed=123, n_predict=64, cache_prompt=false
28/8 raw-completion prompts produced different base vs RWKV-MS continuations.1Prompt: Hi
2Base: ! I'm excited to chat with you. I'm here to help you with whatever you need. Just let me know how I can assist you today!
3RWKV-MS: ! I'm excited to chat with you. What's on your mind today?reports/. That suite mostly
shows semantic equivalence because both endpoints receive the same conversation
history in each request. It should not be read as proof of a memory advantage.1base_model: google/gemma-4-E4B-it
2selected_checkpoint: step-100
3data: tau2 telecom mobile-data rule-planner traces plus format-refresh continuation
4wrapped_layers: 0-5
5delta_heads: q,o
6rank: 8
7alpha: 16
8trainable_memory_params: 797,808
9reported tau2 telecom 20-task screen: 14/20, pass_hat_1=0.70reports/raw_completion_contrast_suite_20260626_225007.mdreports/raw_completion_contrast_suite_20260626_225007.jsonlreports/prompt_suite_compare_20260626_215243.mdreports/prompt_suite_base_20260626_215243.jsonlreports/prompt_suite_rwkv_ms_20260626_215243.jsonlreports/rwkv_ms_runtime_health_20260626_215243.jsonreports/rwkv_ms_memory_sidecar_manifest.jsonreports/gguf_memory_manifest.jsonreports/contrast_trace_20260626_224632.jsonl