Views
No views yet
[!WARNING]🔄 Native Looped vs. Unrolled Architecture Trade-Off:
- Native Looped Binaries (
modelo_qwen3loop_sft_*.gguf): Maximum VRAM and disk compactness (604 MB in Q8_0, 2.8 GB VRAM). Requires our patchedllama.cppbuild or PythonQwen3LoopForCausalLM(trust_remote_code=True).- Unrolled Binaries (
unrolled_modelo_qwen3loop_sft_*.gguf): 100% Universal Compatibility. This version trades off the weight-sharing VRAM benefit (827 MB in Q8_0, ~3.6 GB VRAM), but in exchange runs natively out-of-the-box in standard vanillallama.cpp, Ollama, LM Studio, vLLM, and Hugging Face Transformers WITHOUT requiring any C++ patches or custom Python scripts!
<think>, sharp answers), use the following empirically validated parameters:| Parameter | Recommended Value | Description |
|---|---|---|
| Temperature | 0.60 (or 0.0 for code/math) | Balanced creativity and precision |
| Top-K | 40 | Filters extreme long-tail tokens |
| Top-P | 0.95 | Nucleus sampling threshold |
| Repeat Penalty | 1.12 | Prevents loop degradation in small architectures |
| Repeat Last N | 128 | Repetition penalty lookback buffer |
| Context Window | 32,768 tokens | Native context length |
| Execution Stage / Layer | Before Calibration (Old) | Hotfix 0.2 (Step 75) | Status |
|---|---|---|---|
| Prefix Exit (L06) | 29.10% | 5.20% | Factual Anchoring |
| Loop Pass 1 Exit (L20_p1) | 35.70% | 100.00% | Loop Convergence |
| Loop Pass 2 Exit (L20_p2) ⭐ | 67.20% | 99.80% 🟢 | Peak Recursive Reasoning |
| Suffix Entry (L21) | 44.10% | 99.61% 🟢 | Smooth Thought Projection |
| Suffix Mid (L24) | 27.10% | 96.88% 🟢 | Zero Noise |
| Suffix Final Exit (L27) 🛑 | 21.90% 🛑 | 90.62% 🏆 | Dissipation Eliminated (+68.7%) |
79.38% (77/97) (Net +6.18% gain over DARE-TIES)100.00% (8/8) 🟢100.00% (7/7) 🟢100.00% (7/7) 🟢100.00% (8/8) 🟢90.00% (9/10) 🟢90.00% (9/10) 🟢80.00% (8/10) 🟢80.00% (8/10) 🟢70.00% (7/10) 🟢127.9 - 152.7 tokens/s (NVIDIA RTX 3060 CUDA)unrolled_modelo_qwen3loop_sft_q8_0.gguf (827 MB): Standard Q8_0 GGUF. Works in vanilla llama.cpp and Ollama.unrolled_modelo_qwen3loop_sft_f16.gguf (1.63 GB): Standard FP16 GGUF.modelo_qwen3loop_sft_q8_0.gguf (604 MB): Ultra-compact looped Q8_0 GGUF.modelo_qwen3loop_sft_f16.gguf (1.19 GB): Ultra-compact looped FP16 GGUF.model.safetensors / config.json / tokenizer.json: PyTorch / Transformers model.hotfix0.2.txt: Technical release notes and layerwise probing logs.